Reinforcement Learning for Finite-Horizon H∞ Tracking Control of Unknown Discrete Linear Time-Varying System
针对未知离散线性时变系统,提出了两种强化学习方法(策略迭代和Q学习)来解决有限时域H∞跟踪问题,其中Q学习无需系统模型即可获得控制器,并通过仿真验证了有效性。
This article considers the finite-horizon H<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$_{\infty }$ </tex-math></inline-formula> tracking problem for a class of discrete linear time-varying systems. Two reinforcement learning (RL) methods—policy iteration (PI) and Q-learning—are proposed to solve this problem. The latter can obtain the H<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$_{\infty }$ </tex-math></inline-formula> controller without system dynamics. In the field of RL control, most studies focus on infinite-horizon control and time-invariant systems, and few studies have investigated finite-horizon control or time-varying systems. In contrast to infinite-horizon H<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$_{\infty }$ </tex-math></inline-formula> tracking control, finite-horizon H<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$_{\infty }$ </tex-math></inline-formula> tracking control involves a time-varying value function. While this introduces challenges, it empowers the algorithm to effectively handle time-varying problems. Within the finite-horizon framework, the value function is bounded, allowing the removal of the discount factor, thereby enhancing control performance. Additionally, there is no longer a need for an admissible control law for initialization, providing the proposed algorithms with the combined advantages of both PI and value iteration (VI). Two simulation examples are used to verify the effectiveness of the proposed algorithms.