Roles of pre-training in deep neural networks from information theoretical perspective

Roles of pre-training in deep neural networks from information theoretical perspective
复制标题

DOI:
10.1016/j.neucom.2016.12.083
复制
发表时间:
2017-07-26
期刊:
影响因子:
6
通讯作者:
Ikeda, Kazushi
Ikeda, Kazushi
中科院分区:
计算机科学2区
文献类型:
--
作者:
Furusho, Yasutaka;Kubo, Takatomi;Ikeda, Kazushi

文献摘要

被引文献

相似文献

虽然深度学习在模式识别和机器学习方面表现出很高的性能,但原因仍然不清楚。为了解决这个问题,我们计算了隐藏层中表示的信息理论变量,并分析了它们与性能的关系。我们发现,熵和互信息,这两个都以不同的方式减少,随着层的加深,与微调后的泛化错误。这表明,信息理论变量可能是确定深度学习层数的标准,而无需进行需要高计算负载的微调。(C)2017爱思唯尔B.V.保留所有权利。
Although deep learning shows high performance in pattern recognition and machine learning, the reasons remain unclarified. To tackle this problem, we calculated the information theoretical variables of the representations in the hidden layers and analyzed their relationship to the performance. We found that entropy and mutual information, both of which decrease in a different way as the layer deepens, are related to the generalization errors after fine-tuning. This suggests that the information theoretical variables might be a criterion for determining the number of layers in deep learning without fine-tuning that requires high computational loads. (C) 2017 Elsevier B.V. All rights reserved.