使用折扣逆强化学习的线性连续时间系统输出反馈控制

Output-Feedback Control of Linear Continuous-Time Systems Using Discounted Inverse Reinforcement Learning

IEEE Transactions on Cybernetics · 2026
被引 1 · 同刊同年前 4%
ABS 3

中文导读

提出一种折扣逆强化学习算法,用于部分可观测的连续时间线性系统,仅利用输入输出数据恢复专家控制策略,并证明收敛性和计算效率优势。

Abstract

This article proposes a novel discounted inverse reinforcement learning (DIRL) algorithm for linear quadratic (LQ) control of unknown continuous-time (CT) systems with partially observable states and an unknown discounted value function. Existing DIRL methods predominantly rely on full-state feedback, limiting their applicability to practical scenarios where only input-output data are available. To this end, a state reconstruction method is designed for the system controlled by an expert using the measured desired output. Based on this, a model-free output-feedback (OPFB) DIRL algorithm is presented to iteratively solve the unknown value function and the corresponding optimal OPFB control policy equivalent to the expert control policy. The convergence of the proposed algorithm and the nonuniqueness of solutions are rigorously analyzed. Finally, comprehensive simulations reveal the effectiveness of the proposed algorithm in recovering the expert control policy and its superior computational efficiency compared to state-of-the-art (SOTA) methods.

控制理论强化学习线性系统最优控制