Observer-Based Multi-Agent Reinforcement Learning for Pursuit-Evasion Game With Multiple Unknown Uncertainties
Yangyang Liu, Chun Liu, Yizhen Meng, Bin Jiang, Xiaofan Wang
Shanghai University China Aerospace Science and Technology Corporation Nanjing University of Aeronautics and Astronautics Shanghai Institute of Technology
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
This paper aims to investigate the challenging problem of a multi-agent game with multiple pursuers and a single evader in an environment with multiple unknown uncertainties. A coupled approach combining decentralized observers and reinforcement learning (RL) controllers is proposed to deal with this scenario. Firstly, decentralized observers driven by auxiliary control laws are introduced to estimate the states of uncertain systems, with their best responses obtained through the adaptive dynamic programming (ADP) method. The estimated states, which reflect the actual states of the pursuers’ systems, are concurrently transmitted to the RL controllers. Subsequently, the controllers are trained with observer-based heterogeneous-agent proximal policy optimization (OHAPPO) algorithm, in which a novel global multi-function cost is designed. The algorithm utilizes the advantage decomposition for policy updates in the way of credit assignment, resulting in more stable and efficient updates compared to traditional value decomposition. Moreover, to further enhance the performance of both observers and controllers, a sequential game is established between them, where observers’ policies are influenced by controllers’ optimal control and vice versa. Finally, the simulation results verify the effectiveness of the designed OHAPPO algorithm in the pursuit-evasion game.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
工程Guidance and Control Systems
Evacuation and Crowd Dynamics
参考文献 35
此处列出前 3 条
引用本文 8
按被引量排序,此处列出前 3 条