组合优化中的学习:探索什么与如何探索

Learning in Combinatorial Optimization: What and How to Explore

Operations Research · 2020
被引 12
FT 50UTD 24ABS 4★

中文导读

研究了组合多臂老虎机中的探索与利用权衡,提出下界问题(LBP)和最优性覆盖问题(OCP)策略,在渐近性能上达到理论极限,数值实验显示OCP策略效果显著优于传统方法。

Abstract

When moving from the traditional to combinatorial multiarmed bandit setting, addressing the classical exploration versus exploitation trade-off is a challenging task. In “Learning in Combinatorial Optimization: What and How to Explore,” Modaresi, Sauré, and Vielma show that the combinatorial setting has salient features that distinguish it from the traditional bandit. In particular, combinatorial structure induces correlation between cost of different solutions, thus raising the questions of what parameters to estimate and how to collect and combine information. The authors answer such questions by developing a novel optimization problem called the lower-bound problem (LBP). They establish a fundamental limit on asymptotic performance of any admissible policy and propose near-optimal LBP-based policies. Because LBP is likely intractable in practice, they propose policies that instead solve a proxy for LBP, which they call the optimality cover problem (OCP). They provide strong evidence of practical tractability of OCP and illustrate the markedly superior performance of OCP-based policies numerically.

组合优化多臂老虎机机器学习运筹学