A Class of Bayesian Models for Optimal Exploration
研究了如何最优地探索多个地点以最大化发现收益,将问题建模为贝叶斯序贯决策问题,并发现索引策略最优,其中每个地点未发现物体数量的先验分布尾部行为起关键作用。
SUMMARY Each of several locations contains an unknown number of objects of value. A single search of a location will discover some of these objects which are then removed. Discoveries yield rewards. A ‘distribution of effort’ problem is posed concerning how to explore the locations optimally—i.e. to maximize the expected return from discoveries made. This is formulated as a Bayes sequential decision problem for which index policies are optimal. A natural simple case is one in which, conditionally on the number of undiscovered objects at a location N, the number of discoveries made in a single search is binomial B(N, p) where p is a detection rate. For this case, we can gain considerable insight into how model structure relates to policy structure. The tail behaviour of the priors for the number of objects at each location plays an important role.