Neural self-compressor: Collective interpretation by compressing multi-layered neural networks into non-layered networks

Neural self-compressor: Collective interpretation by compressing multi-layered neural networks into non-layered networks
复制标题

DOI:
10.1016/j.neucom.2018.09.036
复制
发表时间:
2019-01
期刊:
影响因子:
6
通讯作者:
R. Kamimura
R. Kamimura
中科院分区:
计算机科学2区
文献类型:
--
作者:
R. Kamimura

文献摘要

相似文献

本文提出了一种称为“神经自压缩器”的新方法,将多层神经网络压缩成尽可能简单的神经网络(即没有隐层),以帮助解释输入和输出之间的关系。尽管神经网络在改进泛化方面取得了巨大的成功,但随着隐含层的数量及其对应的连接权变得越来越大,内部表示的解释成为一个严重的问题。为了克服这一解释问题,我们引入了一种方法,将多层神经网络压缩为没有隐含层的神经网络。此外,该方法通过最大化输入和输出之间的互信息来尽可能简化纠缠权重。这样,最终连接权重可以像Logistic回归分析一样容易地解释。该方法被应用于四个数据集:对称数据集、卵巢癌数据集、餐馆数据集和信用卡持卡人默认数据集。在第一组对称数据集中,我们试图解释本方法如何能够直观地产生可解释的输出。在所有其他情况下,我们成功地在互信息最大化的帮助下将多层神经网络压缩成最简单的形式。此外,通过解除输出的相关性,我们能够将接近回归系数的连接权重转换为具有更明确特征的连接权重。
The present paper proposes a new method called “neural self-compressors” to compress multi-layered neural networks into the simplest possible ones (i.e., without hidden layers) to aid in the interpretation of relations between inputs and outputs. Though neural networks have shown great success in improving generalization, the interpretation of internal representations becomes a serious problem as the number of hidden layers and their corresponding connection weights becomes larger and larger. To overcome this interpretation problem, we introduce a method that compresses multi-layered neural networks into ones without hidden layers. In addition, this method simplifies entangled weights as much as possible by maximizing mutual information between inputs and outputs. In this way, final connection weights can be interpreted as easily as by the logistic regression analysis. The method was applied to four data sets: a symmetric data set, ovarian cancer data set, restaurant data set, and credit card holders’ default data set. In the first set, the symmetric data set, we tried to explain how the present method could produce interpretable outputs intuitively. In all the other cases, we succeeded in compressing multi-layered neural networks into their simplest forms with the help of mutual information maximization. In addition, by de-correlating outputs, we were able to transform connection weights from those close to the regression coefficients to ones with more explicit features.