Iterative Q -Learning Design for Zero-Sum Games With Evolving Policies
针对动态未知的零和博弈,提出基于值迭代的Q学习框架,分析收敛性和稳定性,并设计两种在线演化控制算法,通过物理实例验证效果。
This article aims to achieve data-based online evolving control for zero-sum games with unknown dynamics. First of all, the value-iteration-based <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Q</i>-learning framework is established. Relevant properties of the iterative <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Q</i>-learning framework are analyzed, including the convergence and monotonicity. Then, the stability property is investigated and the online data is employed for off-policy learning. More importantly, two effective algorithms are designed to achieve online evolving control. In one algorithm, the monotonically nondecreasing <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Q</i>-learning sequence requires the admissible criterion to guarantee the stability with the simple <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Q</i>-function initialization. In another algorithm, the monotonically nonincreasing <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Q</i>-function sequence can ensure the stability without the admissible criterion, but it requires an elaborate initial <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Q</i>-function. In the end, by including two examples of real physical backgrounds, the excellent performance of online evolving control is exhibited with the given algorithms.