基于强化学习的非线性追逃博弈反欺骗方法:不完全与不对称信息

Reinforcement-Learning-Based Counter Deception for Nonlinear Pursuit–Evasion Game With Incomplete and Asymmetric Information

IEEE Transactions on Systems, Man, and Cybernetics: Systems · 2025
被引 3
ABS 3

中文导读

研究了追捕者在信息不完全且不对称时,如何用强化学习应对目标故意隐藏偏好的欺骗行为,提出结合神经网络和卡尔曼滤波的策略,并证明其稳定性。

Abstract

In this article, we investigate the problem of capturing a noncooperative target with deception behavior using reinforcement learning (RL) under incomplete information. The pursuer copes not only with its maneuverability constraint but also with the target’s deception behavior, in which the target deliberately conceals its private preference information. The target capture game involving deception behavior is formulated as a nonlinear differential game framework where the information structure is incomplete and asymmetric. The solution to this differential game is proposed based on an RL policy that incorporates critic, actor, and virtual actor neural networks (NNs), when taking into consideration the maneuverability constraint and information structure of the pursuer. Moreover, the states of the constrained adversarial system and the weight errors are proven to be ultimately uniformly bounded (UUB). To counter the deception of the target, we adopt unscented Kalman filter (UKF) to obtain the target intention on energy preference, and integrate it into the pursuer strategy. The feasibility of the proposed strategy and its superiority are verified through comparisons with recent works.

强化学习博弈论非线性系统欺骗对抗不完全信息