SUBTLEX-CH: Chinese Word and Character Frequencies Based on Film Subtitles
Qing Cai, Marc Brysbaert
HOGENT University of Applied Sciences and Arts Ghent University Hospital Ghent University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
BACKGROUND: Word frequency is the most important variable in language research. However, despite the growing interest in the Chinese language, there are only a few sources of word frequency measures available to researchers, and the quality is less than what researchers in other languages are used to. METHODOLOGY: Following recent work by New, Brysbaert, and colleagues in English, French and Dutch, we assembled a database of word and character frequencies based on a corpus of film and television subtitles (46.8 million characters, 33.5 million words). In line with what has been found in the other languages, the new word and character frequencies explain significantly more of the variance in Chinese word naming and lexical decision performance than measures based on written texts. CONCLUSIONS: Our results confirm that word frequencies based on subtitles are a good estimate of daily language exposure and capture much of the variance in word processing efficiency. In addition, our database is the first to include information about the contextual diversity of the words and to provide good frequency estimates for multi-character words and the different syntactic roles in which the words are used. The word frequencies are freely available for research purposes.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIAuthorship Attribution and Profiling
Text Readability and Simplification · Natural Language Processing Techniques
参考文献 28
此处列出前 3 条
引用本文 796
按被引量排序,此处列出前 3 条