A visual language model enabling intelligent nanomaterial scanning electron micrograph annotation
Yong-Zhu Cai, Hong Wang
Shanghai Jiao Tong University Chinese National Human Genome Center at Shanghai Shanghai Advanced Research Institute
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
contrastive learning on SEM image-text pairs extracted from the literature. SEM-VLM demonstrates superior cross-modal retrieval performance over the general-domain model Contrastive Language-Image Pretraining (CLIP) and random baselines with Recall@10 and Recall@50 metrics, and keyword searches show its robust capability to retrieve relevant images. SEM-VLM also achieves high accuracy in zero-shot classification through ensemble vision-language alignment, outperforming CLIP. In few-shot settings, SEM-VLM with 2.1% training labels exhibits superior performance compared with the fully supervised model (EMCNet: Graph-Nets for Electron Micrograph Classification). Activation mapping analysis reveals precise localization of critical nanoscale features (particles, holes, and probe tips), providing more interpretable results than conventional approaches while maintaining operational reliability. This multimodal framework reduces labeled dataset dependency by orders of magnitude and enables automated high-precision classification.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
材料 / 化学Machine Learning in Materials Science
Corrosion Behavior and Inhibition · Machine Learning in Bioinformatics
参考文献 51
此处列出前 3 条
引用本文 1
按被引量排序,此处列出前 3 条