CMIX: Causal Value Decomposition for Cooperative Multi-Agent Reinforcement Learning
Dunqi Yao, Chuxiong Sun, Kai Li, Kaijie Zhou, Hanyu Li, Rui Wang, Lixiang Liu
University of Chinese Academy of Sciences Chinese Academy of Sciences Institute of Software China Aerospace Science and Industry Corporation (China)
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Value decomposition plays a pivotal role in ensuring effective credit assignment within Multi-Agent Reinforcement Learning (MARL), particularly in cooperative multi-agent tasks where agents are limited to accessing team rewards only. However, existing methods treat the mixing network as a black box, implicitly assuming that neural networks can autonomously extract important information and achieve rational credit assignment during policy learning. This approach not only lacks interpretability but may also prove inefficient in complex scenarios. To enhance the interpretability and rationality of value decomposition, we propose an innovative approach called “Causal Value Decomposition”(CMIX). CMIX employs causal inference-based models, introducing a set of metrics beyond environmental rewards to enhance robustness and model interpretability. Specifically, CMIX establishes intricate relational structures among agents in complex environments and leverages causal relationships between agents and their surroundings to address the credit assignment challenge in MARL. By employing do-calculus, CMIX accurately measures the impact of each agent's actions on environmental states, precisely determining their contribution to the collective reward. This approach not only enhances the interpretability of existing black-box models but also improves the accuracy of credit assignment in multi-agent systems. Moreover, CMIX exhibits high scalability and complements existing value decomposition techniques. Its effectiveness and scalability have been rigorously tested across various settings, including MPE, LBF, and SMAC environments.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIReinforcement Learning in Robotics
参考文献 21
此处列出前 3 条
引用本文 2
按被引量排序,此处列出前 3 条