Assessing the Performance of 8 AI Chatbots in Bibliographic Reference Retrieval: Grok and DeepSeek Outperform ChatGPT, but None are Entirely Accurate
Álvaro Cabezas-Clavijo, Pavel Sidorenko-Bautista
Universidad Internacional De La Rioja
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Purpose This study evaluates the reliability of eight generative artificial intelligence chatbots—including ChatGPT, Claude, Gemini, and DeepSeek—when functioning as autonomous agents for academic bibliographic generation, specifically assessing their accuracy within a university research framework. Design/methodology/approach Using a standardized prompting methodology, 400 references were generated and analyzed across five core knowledge areas: Health, Engineering, Experimental Sciences, Social Sciences, and Humanities. Each agent’s output was rigorously audited against five formal criteria (authorship, year, title, source, and location) and categorized by error frequency and document type. Findings Results indicate a significant reliability gap, with only 26.5 % of references entirely accurate and nearly 40 % flawed or fabricated; while Grok and DeepSeek avoided hallucinations, Copilot, Perplexity, and Claude showed the highest failure rates, particularly when generating journal article citations. Research limitations The study focuses on the free versions of these AI agents, so results may vary with paid models or future architectural updates that integrate real-time web browsing more effectively. Practical implications These findings underscore the critical risks of uncritical reliance on AI agents for academic tasks, highlighting an urgent need for enhanced information literacy and the development of specialized critical thinking skills to navigate AI-mediated research. Originality/value This original and unpublished research provides a pioneering comparative analysis of multiple AI agents as research intermediaries, revealing structural limitations in their generative logic and offering a unique benchmark for the reliability of AI-driven bibliographic data in higher education.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
生物医学Artificial Intelligence in Healthcare and Education
AI in Service Interactions · Topic Modeling
参考文献 44
此处列出前 3 条
引用本文 2
按被引量排序,此处列出前 3 条