Adaptive heterogeneous multi-agent debate for enhanced educational and factual reasoning in large language models
Yan Zhou, Yanguang Chen
South China Agricultural University Shanghai University of Finance and Economics
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Large language models (LLMs) have achieved impressive results in complex reasoning and knowledge tasks, yet they often struggle with factual accuracy and logical consistency. Prior works have improved LLM performance using prompt-based techniques (e.g. chain-of-thought prompting and self-consistency) and post-hoc self-refinement, but these typically operate on a single model instance. Recently, multi-agent debate frameworks have emerged as a complementary approach, wherein multiple LLM agents propose answers and critique each other’s reasoning to reach consensus. Such a “society of minds” approach has been shown to significantly improve mathematical reasoning and reduce factual hallucinations. However, existing debate methods use homogeneous agents with simple majority voting, limiting their effectiveness. In this work, we propose Adaptive Heterogeneous Multi-Agent Debate (A-HMAD), a novel framework that extends multi-agent debate with (i) diverse specialized agents, (ii) dynamic debate routing, and (iii) a learned consensus mechanism. Each agent in A-HMAD is assigned a distinct role or expertise (e.g. logical reasoning, factual verification, strategic planning), enabling more comprehensive error-checking and perspective diversity than identical agents. A coordination policy dynamically selects which agents contribute at each round based on the question’s domain and the evolving debate state. To aggregate viewpoints, we introduce a consensus optimizer that learns to weight each agent’s vote according to its reliability and the confidence of its arguments. On six challenging benchmarks – including arithmetic QA, grade-school math (GSM8K), multifact question answering (MMLU), factual biography generation, and chess strategy – our A-HMAD consistently outperforms prior single-model methods and the original multi-agent debate baseline. Notably, A-HMAD achieves 4–6% absolute accuracy gains over standard debate on these tasks, and reduces factual errors by over 30% in biography facts. We provide extensive ablations demonstrating the benefits of agent heterogeneity, additional debate rounds, and the learned consensus module. Our findings suggest that an adaptive, role-diverse debating ensemble can drive significant advances in LLM-based educational reasoning, paving the way for safer, more interpretable, and pedagogically reliable AI systems.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AITopic Modeling
Multimodal Machine Learning Applications · Advanced Graph Neural Networks
参考文献 4
此处列出前 3 条
引用本文 5
按被引量排序,此处列出前 3 条