基于策略迭代自适应动态规划算法的多人离散时间非零和博弈

Discrete-Time Nonzero-Sum Games for Multiplayer Using Policy-Iteration-Based Adaptive Dynamic Programming Algorithms

IEEE Transactions on Cybernetics · 2016
被引 133
ABS 3

中文导读

提出一种策略迭代自适应动态规划方法,用于求解一类离散时间非线性系统的多人非零和博弈问题,通过迭代控制策略保证系统稳定并最小化各玩家的性能指标,并设计了三种执行器-评判器算法。

Abstract

In this paper, we investigate the nonzero-sum games for a class of discrete-time (DT) nonlinear systems by using a novel policy iteration (PI) adaptive dynamic programming (ADP) method. The main idea of our proposed PI scheme is to utilize the iterative ADP algorithm to obtain the iterative control policies, which not only ensure the system to achieve stability but also minimize the performance index function for each player. This paper integrates game theory, optimal control theory, and reinforcement learning technique to formulate and handle the DT nonzero-sum games for multiplayer. First, we design three actor-critic algorithms, an offline one and two online ones, for the PI scheme. Subsequently, neural networks are employed to implement these algorithms and the corresponding stability analysis is also provided via the Lyapunov theory. Finally, a numerical simulation example is presented to demonstrate the effectiveness of our proposed approach.

自适应动态规划非零和博弈最优控制强化学习非线性系统