基于策略梯度强化学习的二阶多智能体系统数据驱动最优二分一致性控制

Data-Driven Optimal Bipartite Consensus Control for Second-Order Multiagent Systems via Policy Gradient Reinforcement Learning

IEEE Transactions on Cybernetics · 2023
被引 43
ABS 3

中文导读

研究了未知二阶离散时间多智能体系统的最优二分一致性控制问题,利用分布式策略梯度强化学习,提出数据驱动控制策略,保证所有智能体位置和速度状态达到二分一致性,并通过异步算法解决节点计算能力差异问题。

Abstract

This article investigates the optimal bipartite consensus control (OBCC) problem for unknown second-order discrete-time multiagent systems (MASs). First, the coopetition network is constructed to describe the cooperative and competitive relationships between agents, and the OBCC problem is proposed by the tracking error and related performance index function. Based on the distributed policy gradient reinforcement learning (RL) theory, a data-driven distributed optimal control strategy is obtained to guarantee the bipartite consensus of all agents' position and velocity states. In addition, the offline data sets ensure the learning efficiency of the system. These data sets are generated by running the system in real time. Besides, the designed algorithm is an asynchronous version, which is essential to solve the challenge caused by the computational ability difference between nodes in MASs. Then, by means of the functional analysis and Lyapunov theory, the stability of the proposed MASs and the convergence of the learning process are analyzed. Furthermore, an actor-critic structure containing two neural networks is used to implement the proposed methods. Finally, a numerical simulation shows the effectiveness and validity of the results.

多智能体系统强化学习最优控制二分一致性数据驱动控制