基于强化学习和可达性分析的有界理性人机交互自适应安全控制

Adaptive Safe Control for Bounded Rational Human–Robot Interaction Based on Reinforcement Learning and Reachability Analysis

IEEE Transactions on Cybernetics · 2026
被引 0
ABS 3

中文导读

针对人机交互中理性差异和感知限制,提出将交互建模为部分可观测有界理性系统,结合认知层次模型与强化学习,并引入自适应低保守后向可达性分析,在真实交通数据上验证了安全、效率和预测精度的提升。

Abstract

Safety, as a fundamental requirement in human-robot interaction, imposes high demands on autonomous decision-making. Existing studies often model such interaction as a perfect rational game with full observability, ignoring variation in rationality and perceptual constraints. To address these limitations, this article formulates the safe human-robot interaction as a partially observable system with bounded rationality and adopts the cognitive hierarchy (CH) model to characterize the interaction process. However, partial observability leads to discontinuous evolution of game states, resulting in a nonstandard differential game and causing difficulties for theoretical analysis. Based on hybrid system theory and the CH model, this work establishes a relationship between bounded rationality policies and Nash equilibrium, and provides interpretability for integrating CH with reinforcement learning. To further enhance safety, an adaptive low-conservative backward reachability analysis is introduced. Simulation experiments and evaluations on real-world traffic data validate the effectiveness, safety, efficiency, robustness, and predictive accuracy of our framework, reducing conservatism by more than 70% while causing negligible increases in collision risk and achieving approximately 18% higher prediction accuracy compared to another nonequilibrium game model, the level- $k$ model. Furthermore, compared with multiagent reinforcement learning and inverse reinforcement learning (IRL), our framework achieves a better balance between safety and efficiency.

人机交互强化学习安全控制博弈论