基于数据的策略梯度自适应动态规划用于最优控制

Policy Gradient Adaptive Dynamic Programming for Data-Based Optimal Control

IEEE Transactions on Cybernetics · 2016
被引 215 · 同刊同年前 10%
ABS 3

中文导读

针对一般离散时间非线性系统的无模型最优控制问题,提出一种基于数据的策略梯度自适应动态规划算法,利用离线与在线数据改进控制策略,并证明其收敛性。

Abstract

The model-free optimal control problem of general discrete-time nonlinear systems is considered in this paper, and a data-based policy gradient adaptive dynamic programming (PGADP) algorithm is developed to design an adaptive optimal controller method. By using offline and online data rather than the mathematical system model, the PGADP algorithm improves control policy with a gradient descent scheme. The convergence of the PGADP algorithm is proved by demonstrating that the constructed Q -function sequence converges to the optimal Q -function. Based on the PGADP algorithm, the adaptive control method is developed with an actor-critic structure and the method of weighted residuals. Its convergence properties are analyzed, where the approximate Q -function converges to its optimum. Computer simulation results demonstrate the effectiveness of the PGADP-based adaptive control method.

自适应动态规划最优控制非线性系统无模型控制梯度下降