哈密顿驱动的自适应动态规划及其近似误差

Hamiltonian-Driven Adaptive Dynamic Programming With Approximation Errors

IEEE Transactions on Cybernetics · 2021
被引 93
ABS 3

中文导读

针对连续时间非线性系统的无限时域最优控制问题,提出一种哈密顿驱动的迭代自适应动态规划算法,考虑策略评估中的近似误差,并给出保证闭环稳定性和收敛性的条件,还提供了无模型扩展。

Abstract

In this article, we consider an iterative adaptive dynamic programming (ADP) algorithm within the Hamiltonian-driven framework to solve the Hamilton-Jacobi-Bellman (HJB) equation for the infinite-horizon optimal control problem in continuous time for nonlinear systems. First, a novel function, "min-Hamiltonian," is defined to capture the fundamental properties of the classical Hamiltonian. It is shown that both the HJB equation and the policy iteration (PI) algorithm can be formulated in terms of the min-Hamiltonian within the Hamiltonian-driven framework. Moreover, we develop an iterative ADP algorithm that takes into consideration the approximation errors during the policy evaluation step. We then derive a sufficient condition on the iterative value gradient to guarantee closed-loop stability of the equilibrium point as well as convergence to the optimal value. A model-free extension based on an off-policy reinforcement learning (RL) technique is also provided. Finally, numerical results illustrate the efficacy of the proposed framework.

最优控制自适应动态规划强化学习非线性系统