Multi-Step Adaptive Attack Agent: A Dynamic Approach for Jailbreaking Large Language Models
Huiyun Jing, Jincheng Wei, Wei Wei, Yingshui Tan, Boren Zheng, Qingsong Yao
China Academy of Information and Communications Technology Alibaba Group (China) Chinese Academy of Sciences Institute of Computing Technology
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Large Language Models (LLMs) have showcased remarkable potential across various domains, especially in text generation. However, their vulnerability to jailbreak attacks presents considerable challenges to secure deployment, as attackers can use carefully crafted prompts to bypass safety measures and generate harmful content. Current jailbreak methods generally suffer from two significant limitations: a restricted strategy space for generating adversarial prompts and insufficient optimization of prompts based on feedback from LLMs. To overcome these challenges, we present Multistep Adaptive Attack Agent (MATA), an approach that employs a game-theoretic interaction between attack model and target model to adaptively execute jailbreak attacks on LLMs. This method enables iterative attempts based on reflection, gradually identifying the optimal jailbreak attack strategy within a complex strategy space. We compared MATA with mainstream methods across multiple open-source and closed-source LLMs, including Llama3.1, GLM4, and GPT4o. The results demonstrate that our approach exceeds existing methods in terms of attack success rate, average number of queries, and prompt diversity, effectively identifying vulnerabilities in LLMs.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIAdversarial Robustness in Machine Learning
Security and Verification in Computing · Topic Modeling
参考文献 6
此处列出前 3 条