基于分层强化学习的自主空战部分联合优化算法

A Partial Joint Optimization Algorithm for Autonomous Air Combat Based on Hierarchical Reinforcement Learning

IEEE Transactions on Cybernetics · 2025
被引 2
ABS 3

中文导读

提出PJOH-TED2框架,通过部分联合优化机制和时间事件双驱动机制,提升自主空战智能体的探索效率和动态响应能力,在仿真中胜率超71%。

Abstract

Designing intelligent game strategies for autonomous air combat has suffered from the vast exploration space, lengthy decision-making process, and sparse rewards. Some existing approaches adopt the hierarchical framework to improve the exploration efficiency. However, in these methods, agents in different layers are typically trained independently and operate at fixed frequencies, which limits their performance and hampers their ability to respond to highly dynamic combat situations. In view of this, we present PJOH-TED2, a partial-joint-optimization-based hierarchical (PJOH) learning framework with a time-event dual-driven (TED2) mechanism, for one-on-one beyond-visual-range (BVR) air combat. Specifically, the PJOH learning framework embeds the partial joint optimization mechanism into hierarchical reinforcement learning (HRL), thus improving the exploration efficiency dramatically while enhancing the integration across hierarchical levels. Moreover, the TED2 mechanism combines the advantages of event-driven and time-driven methods, which promote the dynamic response speed of agents as well as avoid redundant actions. In addition, we evaluated this work through a series of games against the state-of-the-art (SOTA) methods in a high-fidelity air combat simulation environment. The results empirically demonstrate that the proposed approach outperforms four SOTA methods with a win rate of at least 71%. Finally, this approach achieved the 1st place in learning methods in the intelligent air game algorithm challenge (IAGAC) by the Chinese Institute of Command and Control among 43 teams.

强化学习自主空战分层优化智能博弈