SAFT: Safety-Preserving Adaptation via Fine-Tuning Transfer for Large Language Models
Zhiwen Ruan, Yan Yang, Zhuocheng Liang, Yun Chen, Guanhua Chen
Southern University of Science and Technology Shanghai University of Finance and Economics
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Adapting instruction-tuned large language models (LLMs) to downstream domains is increasingly common, yet fine-tuning on imperfect data can erode the safety alignment learned during post-training. Existing safety-preserving fine-tuning methods typically optimize the aligned instruction model directly, which can destabilize refusal behaviors or impose an ''alignment tax'' that limits task adaptation. We propose SAFT (Safety-preserving Adaptation via Fine-tuning Transfer), a safety-preserving adaptation framework that decouples task learning from alignment preservation by learning a safety-guided task update on the paired pretrained base model, rectifying task gradients to avoid conflicting directions with respect to a safety objective, and then transferring the update to the frozen instruction model via parameter-space grafting. Across mathematical reasoning, code generation, and medical question answering on two open-source model families (Llama3.1-8B-Instruct and Gemma3-4B-IT), SAFT improves downstream utility while maintaining low harmfulness under a unified evaluation protocol, and achieves better safety and utility trade-offs than nine baselines. Warning: This paper contains unfiltered content generated by LLMs that may be offensive to readers.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
暂无主题数据
参考文献 27
此处列出前 3 条