带学习机制的均值场博弈中的熵正则化

Entropy Regularization for Mean Field Games with Learning

Mathematics of Operations Research · 2022
被引 51 · 同刊同年前 3%
ABS 3

中文导读

研究了熵正则化对有限时间均值场博弈中学习的影响,理论证明其能产生时间依赖策略并加速收敛到博弈均衡,进而提出一种带探索的策略梯度算法。

Abstract

Entropy regularization has been extensively adopted to improve the efficiency, the stability, and the convergence of algorithms in reinforcement learning. This paper analyzes both quantitatively and qualitatively the impact of entropy regularization for mean field games (MFGs) with learning in a finite time horizon. Our study provides a theoretical justification that entropy regularization yields time-dependent policies and, furthermore, helps stabilizing and accelerating convergence to the game equilibrium. In addition, this study leads to a policy-gradient algorithm with exploration in MFG. With this algorithm, agents are able to learn the optimal exploration scheduling, with stable and fast convergence to the game equilibrium.

强化学习均值场博弈熵正则化策略梯度算法