基于数据驱动同伦强化学习的马尔可夫跳变非线性系统自适应最优控制

Data-Driven Homotopic Reinforcement Learning-Based Adaptive Optimal Control for Markov Jump Nonlinear Systems

IEEE Transactions on Systems, Man, and Cybernetics: Systems · 2025
被引 12 · 同刊同年前 5%
ABS 3

中文导读

针对连续时间非线性马尔可夫跳变系统,提出一种数据驱动同伦强化学习算法,无需系统模型和初始稳定策略,仅用样本数据即可设计自适应最优控制器。

Abstract

This article investigates the optimal control problem for a class of continuous-time nonlinear Markov jump systems (CTNMJSs), in which an adaptive optimal control policy is developed based on the Takagi–Sugeno (T–S) fuzzy approximation and reinforcement learning (RL) technique. Especially, the original nonlinear system model is first represented in terms of fuzzy rules, and the optimal control problem is transformed into a fuzzy controller design problem for a linear Markov jump fuzzy system without knowledge of the system matrices and input matrices. The policy iteration (PI) algorithm is a powerful RL tool to design adaptive optimal control policy. However, the PI algorithm acquires a stabilizing control policy as its initial policy, which depends extremely on the system dynamics, when system knowledge is unknown, finding an initial stabilizing policy is rather difficult or even impossible. To overcome this shortcoming, in this article, a new off-policy PI-based RL algorithm, i.e., the data-driven homotopic RL (DDHRL), is developed in this article. This DDHRL algorithm improves the traditional PI algorithm, and it is used to obtain the adaptive optimal controller based on the sample data without knowing the system dynamics, and the most significant advantage of the proposed DDHRL is that, by adding a constant sequence, the condition of seeking an initial stabilizing control policy can be avoided. This is in sharp contrast with the traditional PI-based algorithms. The convergence of the DDHRL algorithm is proved, and its feasibility and good performance are validated by simulation examples.

强化学习非线性系统最优控制马尔可夫跳变系统模糊控制