Discovery of CRISPR-Cas12a clades using a large language model
Yuanyuan Feng, Junchao Shi, Zhan‐Wei Li, Yongqian Li, Jiaxi Yang, Shisheng Huang, Jinfang Zheng, Wei Han 等 18 位
Zhejiang Lab Shanghai Jiao Tong University Shanghai Ninth People's Hospital State Key Laboratory of Reproductive Medicine
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
The identification and engineering of CRISPR-Cas systems revolutionized life science. Metagenome contains millions of unknown Cas proteins, which require precise prediction and characterization. Traditional protein mining mainly depends on protein sequence alignments. In this work, we harnessed the capability of the evolutionary scale language model (ESM) to learn the information beyond the sequence. After training with the CRISPR-Cas sequences and their functional annotation, the ESM model can identify the CRISPR-Cas proteins from the annotated genome sequences accurately and robustly without sequence alignment. However, due to the lack of experimental data, the feature prediction is limited by the small sample size. Integrated with machine learning on small size experimental data, the model is able to predict the trans-cleavage activity of novel Cas12a. Furthermore, we discovered 7 novel subtypes of Cas12a proteins with unique organization of CRISPR loci and protein sequences. Notably, structural alignments revealed that Cas1, Cas2, and Cas4 also exhibit 8 subtypes, with the absence of integrase proteins correlating with a reduction in spacer numbers within CRISPR loci. In addition, the Cas12a subtypes displayed distinct 3D foldings, a finding further corroborated by CryoEM analyses that unveiled unique interaction patterns with RNA. Accordingly, these proteins show distinct double-strand and single-strand DNA cleavage preferences and broad PAM recognition. Finally, we established a specific detection strategy for the oncogene SNP without traditional Cas12a PAM. This study shows the great potential of the language model in the novel Cas protein function exploration via gene cluster classification.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
生物医学CRISPR and Genetic Engineering
Machine Learning in Bioinformatics · Genomics and Phylogenetic Studies
参考文献 62
此处列出前 3 条
引用本文 21
按被引量排序,此处列出前 3 条