CANDY: Benchmarking LLMs’ Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
Ruiling Guo, Xinwei Yang, Chen Huang, Tong Zhang, Yong Kai Hu
Sichuan University National University of Singapore Institute of Data Science
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
The effectiveness of large language models (LLMs) to fact-check misinformation remains uncertain, despite their growing use.To this end, we present CANDY, a benchmark designed to systematically evaluate the capabilities and limitations of LLMs in fact-checking Chinese misinformation.Specifically, we curate a carefully annotated dataset of ∼20k instances.Our analysis shows that current LLMs exhibit limitations in generating accurate fact-checking conclusions, even when enhanced with chainof-thought reasoning and few-shot prompting.To understand these limitations, we develop a taxonomy to categorize flawed LLM-generated explanations for their conclusions and identify factual fabrication as the most common failure mode.Although LLMs alone are unreliable for fact-checking, our findings indicate their considerable potential to augment human performance when deployed as assistive tools in scenarios.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
社会科学Misinformation and Its Impacts
Topic Modeling · Computational and Text Analysis Methods
参考文献 0
引用本文 2
按被引量排序,此处列出前 3 条