ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations
Yunze Xiao, Yujia Hu, Kenny Tsu Wei Choo, Roy Ka-Wei Lee
Carnegie Mellon University Qatar Carnegie Mellon University Singapore University of Technology and Design
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Detecting hate speech and offensive language is essential for maintaining a safe and respectful digital environment.This study examines the limitations of state-of-the-art large language models (LLMs) in identifying offensive content within systematically perturbed data, with a focus on Chinese, a language particularly susceptible to such perturbations.We introduce ToxiCloakCN 1 , an enhanced dataset derived from ToxiCN, augmented with homophonic substitutions and emoji transformations, to test the robustness of LLMs against these cloaking perturbations.Our findings reveal that existing models significantly underperform in detecting offensive content when these perturbations are applied.We provide an in-depth analysis of how different types of offensive content are affected by these perturbations and explore the alignment between human and model explanations of offensiveness.Our work highlights the urgent need for more advanced techniques in offensive language detection to combat the evolving tactics used to evade detection mechanisms.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIHate Speech and Cyberbullying Detection
Natural Language Processing Techniques · Interpreting and Communication in Healthcare
参考文献 0
引用本文 12
按被引量排序,此处列出前 3 条