Impulsive maneuver spacecraft pursuit-defense game based on multi-agent reinforcement learning
Fan Shuhui, Xiang Zhang, Liao Wenhe
Nanjing University of Science and Technology
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
To address long-range impulsive orbital pursuit-defense game scenario, characterized by long decision cycles, sparse actions, and misaligned maneuvering times, a multiagent asynchronous hierarchical reinforcement learning (MA-AHRL) method is proposed. The method considers practical constraints including space perturbations, thrust limitations, fuel consumption limitations, and perceptual capabilities. Firstly, a relative orbital dynamics model is established, and the pursuit-defense game environment is defined, which clarifies the success criteria for the tasks of both the pursuer and the defender. Secondly, an actor-critic network structure integrating long short-term memory (LSTM) networks with an adaptive attention mechanism is designed. The LSTM extracts features from historical interaction data to compensate for the critic network's inability to access opponent action information during training. The attention mechanism adaptively adjusts the weights of different time steps, thereby enhancing the ability to extract key state features in the game environment. Then, the action network employs a dual-layer decision structure that integrates dynamics, where the high-level policy outputs the maneuver time interval and the low-level policy outputs impulsive actions, effectively reducing training difficulty and improving training efficiency. Finally, the pursuit-defense game is modeled as a partially observable Markov decision process (POMDP), where the action space, state space, observation space, and reward function are explicitly designed. Simulation comparison demonstrates that MA-AHRL can effectively coordinate the competitive relationship between the pursuer and defender while exhibiting excellent training stability. Testing the decision-making networks, both the pursuer and the defender display strong strategic adaptability and real-time decision-making capabilities, fully validating the effectiveness and reliability of the proposed method in this paper.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
工程Guidance and Control Systems
Reinforcement Learning in Robotics · Space Satellite Systems and Control
参考文献 0
引用本文 1
按被引量排序,此处列出前 3 条