星际航天鲁棒低推力轨迹设计:一种自适应潜在强化学习方法

Robust Low-Thrust Trajectory Design for Interplanetary Spaceflight: An Adaptive Latent Reinforcement Learning Method

IEEE Transactions on Cybernetics · 2026
被引 0
ABS 3

中文导读

针对低推力航天器在状态和观测不确定性下的鲁棒轨迹设计问题,提出一种基于序列潜在变量模型的自适应潜在强化学习方案,通过从潜在变量而非原始观测中推导控制策略来缓解不确定性影响,并在两个交会任务中验证了有效性。

Abstract

This article investigates the problem of robust trajectory design for low-thrust spacecraft subject to state and observation uncertainties. An adaptive latent reinforcement learning (RL) scheme based on sequential latent variable models (SLVMs) is proposed to address this issue. First, an SLVM is employed for the representation learning of uncertain environments and for predicting future observations. Subsequently, by integrating representation learning based on the SLVM with proximal policy optimization (PPO), a stochastic latent PPO (SLPPO) scheme is introduced. Distinct from existing methods, the control policy is derived from learned stochastic latent variables rather than raw uncertain observations, which effectively mitigates the adverse impact of uncertainties on control performance. Furthermore, to enhance training efficiency, an improved dense reward shaping mechanism is designed based on the observation predictions from the SLVM and adaptive techniques. Finally, numerical simulations of two rendezvous missions validate the effectiveness of the proposed approach.

航天器轨迹优化强化学习不确定性控制自适应控制