Fusing finetuned models for better pretraining
Leshem Choshen, Elad Venezian, Noam Slonim, Yoav Katz
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Pretrained models are the standard starting point for training. This approach consistently outperforms the use of a random initialization. However, pretraining is a costly endeavour that few can undertake. In this paper, we create better base models at hardly any cost, by fusing multiple existing fine tuned models into one. Specifically, we fuse by averaging the weights of these models. We show that the fused model results surpass the pretrained model ones. We also show that fusing is often better than intertraining. We find that fusing is less dependent on the target task. Furthermore, weight decay nullifies intertraining effects but not those of fusing.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIMachine Learning and Algorithms
Machine Learning and Data Classification · Domain Adaptation and Few-Shot Learning
参考文献 0
引用本文 12
按被引量排序,此处列出前 3 条