研究论文开放获取
Deduplicating Training Data Makes Language Models Better
Katherine Lee, Daphne Ippolito, A. Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, Nicholas Carlini
Google (United States) Brain (Germany) California University of Pennsylvania
来源Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
年份2022
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
摘要 · 完整
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, Nicholas Carlini. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
逐年被引趋势
75380
21
22
23
24
7525
26
关键指标
288
被引次数 · OpenAlex
36.56
领域内被引倍数
同类平均 = 1
同类平均 = 1
前 0.1%
引用位次
同领域 · 同年份 · 同类型
同领域 · 同年份 · 同类型
58
参考文献
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
论文问答
当前基于摘要回答
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AITopic Modeling
Natural Language Processing Techniques · Data Quality and Management
参考文献 58
Space Efficient Linear Time Construction of Suffix Arrays
被引 225Pang Ko, Srinivas Aluru · Lecture notes in computer science · 2003
One billion word benchmark for measuring progress in statistical language modeling
被引 637Ciprian Chelba, Tomáš Mikolov, Mike Schuster · 2014
Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
被引 2,083Yukun Zhu, Ryan Kiros, Rich Zemel · 2015
此处列出前 3 条
引用本文 288
Survey of Hallucination in Natural Language Generation
被引 4,542Ziwei Ji, Nayeon Lee, Rita Frieske · ACM Computing Surveys · 2022
On the Opportunities and Risks of Foundation Models
被引 2,296Rishi Bommasani, Drew A. Hudson, Ehsan Adeli · arXiv (Cornell University) · 2021
A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
被引 2,103Lei Huang, Weijiang Yu, Weitao Ma · ACM Transactions on Information Systems · 2024
按被引量排序,此处列出前 3 条