延迟响应多臂赌博机的Gittins过程

Gittins Procedures for Bandits with Delayed Responses

Journal of the Royal Statistical Society. Series B: Statistical Methodology · 1988
被引 21
ABS 4

中文导读

本文研究带几何折扣的延迟响应多臂赌博机,证明折扣因子小于1/2或信息库为零时存在动态分配指数,并给出计算指数的方法,对金融、人工智能等领域有用。

Abstract

SUMMARY This paper introduces the multi-armed delayed response bandit with geometric discounting. The existence of dynamic allocation indices is shown when the discount factor is less than 1/2 or when the information bank size is zero. For the multi-armed delayed response bandit, the arm indicated by the dynamic allocation procedure or Gittins procedure is optimal when all information bank sizes are zero. A computational method for calculating indices is presented. The idea is to approximate the optimal strategy using a class of strategies whose worths are easy to calculate.

多臂赌博机动态分配折扣因子最优策略计算算法