面向自动驾驶的长期与短期约束驱动安全强化学习

Long- and Short-Term Constraint-Driven Safe Reinforcement Learning for Autonomous Driving

IEEE Transactions on Systems, Man, and Cybernetics: Systems · 2026
被引 1 · 同刊同年前 7%
ABS 3

中文导读

提出一种结合长期和短期约束的安全强化学习算法,通过拉格朗日乘子优化训练过程,在自动驾驶仿真中提升了成功率并降低了事故成本。

Abstract

Safe reinforcement learning (RL) is developed to handle high-risk decision-making tasks, such as autonomous driving (AD), by constraining expected safety violation costs as a training objective. However, existing safe RL methods only consider the long-term objective but ignore the short-term state safety of exploration in the training process. In addition, it is difficult to achieve a balance between cost and return expectations, leading to deterioration of learning performance. Unlike these methods, we propose a novel algorithm named long-and short-term constraints (LSTCs) for safe RL. The short-term constraint is proposed to enhance the short-term state safety that the vehicle explores, while the long-term constraint enhances the overall safety of the vehicle throughout the decision-making process, both of which are jointly used to enhance vehicle safety in the training process. Furthermore, we develop a safe RL method with dual-constraint optimization based on the Lagrange multiplier to optimize the training process for end-to-end AD, balancing the cost and return expectations. Comprehensive experiments were conducted on the MetaDrive simulator. The experimental results demonstrate that the success rate increases by 13% and the episode cost decreases by 0.26 compared to the best results of the comparative methods, showing that the proposed method has better safety in continuous control tasks and exhibits a higher exploration performance in long-distance decision-making tasks compared to SOTA methods.

自动驾驶安全强化学习约束优化决策控制