基于双模块学习MATD3的无人艇集群自主路径规划

Autonomous Path Planning for USV Swarm Based on Dual-Module Learning MATD3

IEEE Transactions on Cybernetics · 2026
被引 0
ABS 3

中文导读

针对复杂海洋环境下无人艇集群路径规划问题,提出双模块学习框架DML-MATD3,通过运动生成和避碰模块分别设计奖励函数,结合势场法和自适应能耗奖励,实现更快收敛和更优路径。

Abstract

The complex and dynamic maritime environment brings significant uncertainties to autonomous path planning tasks, posing substantial challenges for multiagent reinforcement learning (MARL). To address these challenges, this article analyzes the kinematic and dynamic models of unmanned surface vessels (USVs) as well as the environmental disturbances (e.g., wind, waves, and currents), and then proposes a dual-module learning multiagent twin delayed deep deterministic (DML-MATD3) policy gradient framework for USV swarm path planning based on the realistic physical conditions. The framework establishes a motion-generation module and a collision-avoidance module to enable simpler yet more effective reward designs for learning. Specifically, tailored reward functions are independently designed for each module. A potential field method (PFM) is introduced to provide dense and informative guidance for both motion-generation and collision-avoidance modules. Moreover, an adaptive energy consumption reward (AECR) is integrated into the motion-generation module to improve energy-efficient navigation under environmental disturbances. To further enhance exploration efficiency and reward responsiveness during early training, an Ornstein-Uhlenbeck noise-based action enhancement strategy (OU-AES) is employed. Extensive experiments against seven baseline algorithms demonstrate that the proposed DML-MATD3 consistently achieves faster convergence, improved training stability, shorter path lengths, reduced task execution times, and superior overall performance in complex maritime environments.

无人艇路径规划多智能体强化学习自主导航