基于强化学习和并行复合策略迭代的未知动态非线性系统模糊控制

Reinforcement Learning-Based Fuzzy Control for Nonlinear Systems With Unknown Dynamics via Parallel Composite Policy Iteration Scheme

IEEE Transactions on Cybernetics · 2026
被引 1 · 同刊同年前 4%
ABS 3

中文导读

提出一种并行复合策略迭代算法,解决未知动态非线性系统的模糊控制问题,无需初始稳定控制策略和持续激励条件,通过在线数据与历史数据结合降低数据需求,并在单连杆机械臂和四分之一车主动悬架实验中验证有效性。

Abstract

The problem of reinforcement learning (RL)-based fuzzy control for nonlinear systems with unknown dynamics via parallel composite policy iteration (PCPI) scheme is studied in this article. The main objective of this article is to solve the fuzzy algebraic Riccati equation (FARE), which is inherently complex and cannot be easily solved by traditional mathematical formulas. Policy iteration (PI) and value iteration (VI) algorithms proposed have been widely used to address this problem. However, these algorithms have the disadvantages of an initial stabilizing control policy, the persistent excitation (PE) condition, and huge amounts of data. To effectively alleviate these drawbacks, a novel PCPI algorithm is proposed in this article. Specifically, for each fuzzy subsystem, an adaptive parameter is designed to eliminate the requirement of an initial stabilizing control policy. In addition, an online model-free PCPI algorithm is proposed for the situation where the dynamic information of the fuzzy system is difficult to obtain. By substituting the stored historical data with online data, the PE condition is relaxed to the initial excitation (IE) condition. Concurrently, the corresponding algorithm can be executed independently and concurrently under each fuzzy rule, thereby fully exploiting the available computational resources. Finally, the effectiveness of the algorithms set forth in this article is verified through a single-link robot arm and quarter-car active suspension (QCAS) experiment.

强化学习模糊控制非线性系统自适应控制