SafeLawBench: Towards Safe Alignment of Large Language Models
Chuxue Cao, Zhu Han, Jiaming Ji, Qichao Sun, Zining Zhu, Wu Yinyu, Josef Dai, Yaodong Yang 等 10 位
Hong Kong University of Science and Technology Peking University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
With the growing prevalence of large language models (LLMs), the safety of LLMs has raised significant concerns.However, there is still a lack of definitive standards for evaluating their safety due to the subjective nature of current safety benchmarks.To address this gap, we conducted the first exploration of LLMs' safety evaluation from a legal perspective by proposing the SafeLawBench benchmark.SafeLawBench categorizes safety risks into three levels based on legal standards, providing a systematic and comprehensive framework for evaluation.It comprises 24,860 multi-choice questions and 1,106 open-domain question-answering (QA) tasks.Our evaluation included 2 closed-source LLMs and 18 open-source LLMs using zero-shot and fewshot prompting, highlighting the safety features of each model.We also evaluated the LLMs' safety-related reasoning stability and refusal behavior.Additionally, we found that a majority voting mechanism can enhance model performance.Notably, even leading SOTA models like Claude-3.5-Sonnetand GPT-4o have not exceeded 80.5% accuracy in multi-choice tasks on SafeLawBench, while the average accuracy of 20 LLMs remains at 68.8%.We urge the community to prioritize research on the safety of LLMs.Our dataset and code are available.1
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AITopic Modeling
Natural Language Processing Techniques
参考文献 0
引用本文 4
按被引量排序,此处列出前 3 条