离散时间非线性系统最优控制的值迭代自适应动态规划

Value Iteration Adaptive Dynamic Programming for Optimal Control of Discrete-Time Nonlinear Systems

IEEE Transactions on Cybernetics · 2015
被引 452 · 同刊同年前 1%
ABS 3

中文导读

提出一种值迭代自适应动态规划算法,用于求解离散时间非线性系统的无限时域无折扣最优控制问题,允许任意半正定函数初始化,并保证迭代值函数收敛到最优性能指标。

Abstract

In this paper, a value iteration adaptive dynamic programming (ADP) algorithm is developed to solve infinite horizon undiscounted optimal control problems for discrete-time nonlinear systems. The present value iteration ADP algorithm permits an arbitrary positive semi-definite function to initialize the algorithm. A novel convergence analysis is developed to guarantee that the iterative value function converges to the optimal performance index function. Initialized by different initial functions, it is proven that the iterative value function will be monotonically nonincreasing, monotonically nondecreasing, or nonmonotonic and will converge to the optimum. In this paper, for the first time, the admissibility properties of the iterative control laws are developed for value iteration algorithms. It is emphasized that new termination criteria are established to guarantee the effectiveness of the iterative control laws. Neural networks are used to approximate the iterative value function and compute the iterative control law, respectively, for facilitating the implementation of the iterative ADP algorithm. Finally, two simulation examples are given to illustrate the performance of the present method.

最优控制自适应动态规划非线性系统离散时间系统