均值-方差设定下的多臂老虎机问题

The multi-armed bandit problem under the mean-variance setting

European Journal of Operational Research · 2025
被引 1
ABS 4

中文导读

研究了均值-方差设定下的多臂老虎机问题,放宽了独立臂和有界奖励的假设,提出风险感知下置信界算法,数值模拟显示其优于现有方法。

Abstract

The classical multi-armed bandit problem involves a learner and a collection of arms with unknown reward distributions. At each round, the learner selects an arm and receives new information. The learner faces a tradeoff between exploiting the current information and exploring all arms. The objective is to maximize the expected cumulative reward over all rounds. Such an objective does not involve a risk-reward tradeoff, which is fundamental in many areas of application. In this paper, we build upon Sani et al. (2012)’s extension of the classical problem to a mean–variance setting. We relax their assumptions of independent arms and bounded rewards, and we consider sub-Gaussian arms. We introduce the Risk-Aware Lower Confidence Bound algorithm to solve the problem, and study some of its properties. We perform numerical simulations to demonstrate that, in both independent and dependent scenarios, our approach outperforms the algorithm suggested by Sani et al. (2012).

多臂老虎机均值-方差优化风险感知算法强化学习