DocAgent: An Agentic Framework for Multi-Modal Long-Context Document Understanding
Li Sun, He Liu, Shuyue Jia, Yangfan He, Chenyu You
Boston University University of Minnesota Stony Brook University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Recent advances in large language models (LLMs) have demonstrated significant promise in document understanding and questionanswering.Despite the progress, existing approaches can only process short documents due to limited context length or fail to fully leverage multi-modal information.In this work, we introduce DocAgent, a multi-agent framework for long-context document understanding that imitates the human reading practice.Specifically, we first extract a structured, tree-formatted outline from documents to help agents identify relevant sections efficiently.Further, we develop an interactive reading interface that enables agents to query and retrieve various types of content dynamically.To ensure answer reliability, we introduce a reviewer agent that cross-checks responses using complementary sources and maintains a task-agnostic memory bank to facilitate knowledge sharing across tasks.We evaluate our method on two long-context document understanding benchmarks, where it bridges the gap to human-level performance by surpassing competitive baselines, while maintaining a short context length.Our code is available at https://github.com/lisun-ai/DocAgent.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AITopic Modeling
Text Readability and Simplification · Speech and dialogue systems
参考文献 0
引用本文 3
按被引量排序,此处列出前 3 条