Adaptive Output Synchronization With Designated Convergence Rate of Multiagent Systems Based on Off-Policy Reinforcement Learning
研究了线性离散时间多智能体系统的H∞最优输出同步问题,提出一种离策略强化学习方法,仅利用输入输出数据在线学习同步协议,使系统以指定速度渐近同步。
In this article, an optimal output synchronization solution to the <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$H_{\infty}$</tex-math> </inline-formula> optimization of linear discrete-time (DT) multiagent systems is investigated. Compared with current approaches, the issue of designated convergence rate is handled with system optimality, while less computation cost is required. Specifically, the internal model principle is employed to derive a cooperative regulation problem of DT systems, wherein no explicit solution to output regulation equations is needed for learning. Then, we introduce a convergence rate parameter to construct a group of auxiliary cooperative systems, based on which the zero-sum game in <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$H_{\infty}$</tex-math> </inline-formula> optimization is formulated. The data-efficient off-policy reinforcement learning and output-feedback technique are applied to solve the enhanced Bellman equations with a designated convergence rate. This results in an online optimal synchronization solution learning from only the input–output data along the system trajectories. It is shown that the proposed optimal synchronization protocol achieves asymptotic synchronization for the original systems with the consensus error converging to zero at a designated rate. The effectiveness of the proposed approach is verified by the simulation results.