Exploratory Experiments in Generating Bag of Word Dictionaries Using Word2Vec and ChatGPT
Daniel E. O’Leary, Yangin Yoon, Kevin Moffitt
University of Southern California California Southern University Seoul National University of Science and Technology Rutgers, The State University of New Jersey
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
This paper investigates the use of machine learning as a means of generating “bag of words” using text corpora from accounting applications. We use Word2Vec to generate words/dictionaries, that are “similar” to a seed word that captures a concept. As part of our analysis, we perform several experiments using text from Form 10Ks and earnings calls. We investigate several activities including choice of the seed word(s), choosing word sources (corpora), analysis of resulting word lists, and other concerns. We also examine the notion of “human-in-the-loop” and the roles that a person would need to perform while generating a dictionary. Further, we investigate the impact of using accounting and financial corpuses on the different semantic and syntactic relationships, in contrast to Wikipedia. We then extend the analysis to compare those findings to ChatGPT another source of words and investigate some of the advantages and disadvantages of that approach. Data Availability: Data are available from the authors. JEL Classifications: M4.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AINatural Language Processing Techniques
Lexicography and Language Studies
参考文献 31
此处列出前 3 条
引用本文 1
按被引量排序,此处列出前 3 条