The Impact of Tokenizer Selection in Genomic Language Models
LeAnn M. Lindsey, Nicole L. Pershing, A K M Rubaiyat Reza Habib, Keith Dufault‐Thompson, W. Zac Stephens, Anne J. Blaschke, Xiaofang Jiang, Hari Sundar
National Institutes of Health University of Utah United States National Library of Medicine Tufts University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Genomic language models have recently emerged as a new method to decode, interpret, and generate genetic sequences. Existing genomic language models have utilized various tokenization methods, including character tokenization, overlapping and non-overlapping k-mer tokenization, and byte-pair encoding, a method widely used in natural language models. Genomic sequences differ from natural language because of their low character variability, complex and overlapping features, and inconsistent directionality. These features make sub-word tokenization in genomic language models significantly different from both traditional language models and protein language models. This study explores the impact of tokenization in genomic language models by evaluating their downstream performance on forty-four classification fine-tuning tasks. We also perform a direct comparison of byte pair encoding and character tokenization in Mamba, a state-space model. Our results indicate that character tokenization outperforms sub-word tokenization methods on tasks that rely on nucleotide level resolution, such as splice site prediction and promoter detection. While byte-pair tokenization had stronger performance on the SARS-CoV-2 variant classification task, we observed limited statistically significant differences between tokenization methods on the remaining downstream tasks.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
生物医学Genomics and Phylogenetic Studies
Machine Learning in Bioinformatics · RNA and protein synthesis mechanisms
参考文献 28
此处列出前 3 条
引用本文 7
按被引量排序,此处列出前 3 条