Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
Dan Hendrycks, Thomas G. Dietterich
University of California, Berkeley Oregon State University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Convolutional neural networks achieve near-perfect accuracy on curated plant disease image datasets, yetperformance falls sharply when the same models encounter real field imagery. Existing work documents thislab-to-field gap and catalogues its causes qualitatively without quantifying how much each capture conditioncontributes. This paper measures that contribution directly. An EfficientNet-B0 classifier was trained bytransfer learning from ImageNet on a three-class tomato subset of PlantVillage (Early Blight, Late Blight,Healthy), then frozen. The frozen model was evaluated once against a clean 677-image test set and againagainst twelve synthetically degraded versions of that same test set, comprising four degradation conditions(brightness and contrast variation, Gaussian blur, Gaussian noise, and resolution reduction) at three severitylevels each. Freezing the weights isolates image condition as the sole experimental variable. Clean accuracywas 94.98%, with a macro-averaged F1-score of 94.55%. Robustness varied by condition: accuracy remained at91.43% under severe Gaussian noise but fell to 49.04% under severe Gaussian blur, a decline of 45.94percentage points. Blur also produced a cliff-shaped rather than gradual decline, offering no early warningbefore failure. Across the three most damaging conditions macro recall fell faster than macro precision,indicating that degradation drives the model toward missed disease rather than false alarms, the costlier errorin agriculture. A parallel model trained from random initialization showed that the transfer-learningadvantage, 19.65 points on clean data, narrowed under noise and reversed entirely under severe blur andsevere brightness and contrast degradation. Accuracy claims made on curated data therefore overstate fieldreliability, and architecture choice should reflect the capture conditions a deployment will face.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIAdversarial Robustness in Machine Learning
参考文献 0
引用本文 306
按被引量排序,此处列出前 3 条