Gittins Procedures for Bandits with Delayed Responses
本文研究带几何折扣的延迟响应多臂赌博机,证明折扣因子小于1/2或信息库为零时存在动态分配指数,并给出计算指数的方法,对金融、人工智能等领域有用。
SUMMARY This paper introduces the multi-armed delayed response bandit with geometric discounting. The existence of dynamic allocation indices is shown when the discount factor is less than 1/2 or when the information bank size is zero. For the multi-armed delayed response bandit, the arm indicated by the dynamic allocation procedure or Gittins procedure is optimal when all information bank sizes are zero. A computational method for calculating indices is presented. The idea is to approximate the optimal strategy using a class of strategies whose worths are easy to calculate.