基于逆最优控制的跟踪控制中的逆强化学习

Inverse Reinforcement Learning in Tracking Control Based on Inverse Optimal Control

IEEE Transactions on Cybernetics · 2021
被引 130 · 同刊同年前 6%
ABS 3

中文导读

提出一种新的逆强化学习算法,通过结合最优控制更新、梯度下降修正和逆最优控制更新,学习未知的性能目标函数,用于跟踪控制,并分析了奖励权重的非唯一性。

Abstract

This article provides a novel inverse reinforcement learning (RL) algorithm that learns an unknown performance objective function for tracking control. The algorithm combines three steps: 1) an optimal control update; 2) a gradient descent correction step; and 3) an inverse optimal control (IOC) update. The new algorithm clarifies the relation between inverse RL and IOC. It is shown that the reward weight of an unknown performance objective that generates a target control policy may not be unique. We characterize the set of all weights that generate the same target control policy. We develop a model-based algorithm and, further, two model-free algorithms for systems with unknown model information. Finally, simulation experiments are presented to show the effectiveness of the proposed algorithms.

强化学习最优控制跟踪控制逆问题机器学习