Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution
Shuo Shao, Yiming Li, Hongwei Yao, Yiling He, Zhan Qin, Kui Ren
Zhejiang University Nanyang Technological University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Ownership verification is currently the most critical and widely adopted post-hoc method to safeguard model copyright.In general, model owners exploit it to identify whether a given suspicious third-party model is stolen from them by examining whether it has particular properties 'inherited' from their released models.Currently, backdoor-based model watermarks are the primary and cutting-edge methods to implant such properties in the released models.However, backdoor-based methods have two fatal drawbacks, including harmfulness and ambiguity.The former indicates that they introduce maliciously controllable misclassification behaviors (i.e., backdoor) to the watermarked released models.The latter denotes that malicious users can easily pass the verification by finding other misclassified samples, leading to ownership ambiguity.In this paper, we argue that both limitations stem from the 'zero-bit' nature of existing watermarking schemes, where they exploit the status (i.e., misclassified) of predictions for verification.Motivated by this understanding, we design a new watermarking paradigm, i.e., Explanation as a Watermark (EaaW), that implants verification behaviors into the explanation of feature attribution instead of model predictions.Specifically, EaaW embeds a 'multi-bit' watermark into the feature attribution explanation of specific trigger samples without changing the original prediction.We correspondingly design the watermark embedding and extraction algorithms inspired by explainable artificial intelligence.In particular, our approach can be used for different tasks (e.g., image classification and text generation).Extensive experiments verify the effectiveness and harmlessness of our EaaW and its resistance to potential attacks.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIAdversarial Robustness in Machine Learning
Advanced Steganography and Watermarking Techniques · Physical Unclonable Functions (PUFs) and Hardware Security
参考文献 0
引用本文 30
按被引量排序,此处列出前 3 条