Deep learning and the information bottleneck principle
Naftali Tishby, Noga Zaslavsky
Hebrew University of Jerusalem
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Deep Neural Networks (DNNs) are analyzed via the theoretical framework of the information bottleneck (IB) principle. We first show that any DNN can be quantified by the mutual information between the layers and the input and output variables. Using this representation we can calculate the optimal information theoretic limits of the DNN and obtain finite sample generalization bounds. The advantage of getting closer to the theoretical limit is quantifiable both by the generalization bound and by the network's simplicity. We argue that both the optimal architecture, number of layers and features/connections at each layer, are related to the bifurcation points of the information bottleneck tradeoff, namely, relevant compression of the input layer with respect to the output layer. The hierarchical representations at the layered network naturally correspond to the structural phase transitions along the information curve. We believe that this new insight can lead to new optimality bounds and deep learning algorithms.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
物理Statistical Mechanics and Entropy
Neural Networks and Applications · Gaussian Processes and Bayesian Inference
参考文献 17
此处列出前 3 条
引用本文 1,467
按被引量排序,此处列出前 3 条