Information-Theoretic Competitive Learning with Inverse Euclidean Distance Output Units

Information-Theoretic Competitive Learning with Inverse Euclidean Distance Output Units
复制标题

具有反欧几里德距离输出单元的信息论竞争学习

DOI:
10.1023/b:nepl.0000011136.78760.22
复制
发表时间:
2003
影响因子:
3.1
通讯作者:
R. Kamimura
R. Kamimura
中科院分区:
计算机科学4区
文献类型:
--
作者:
R. Kamimura

文献摘要

被引文献

相似文献

本文提出了一种新的信息论竞争学习方法。我们首先在单层网络中构建了一种学习方法,然后将其推广到有监督的多层网络中。竞争单元输出由输入模式与连接权值之间的欧氏距离的倒数计算。距离越小,竞争单位产出越强。在实现竞争时,既不使用赢者通吃算法,也不使用横向抑制算法。相反,新方法是基于输入模式和竞争单位之间的相互信息最大化。在互信息最大化中,竞争单位的熵尽可能地增加。这意味着在我们的框架中必须平等地使用所有竞争单位。因此,不会产生未充分利用的神经元或死亡神经元。当使用多层网络时,通过统一信息最大化和最小化,可以提高网络的抗噪性能。我们将单层网络的方法应用于一个简单的人工数据问题和一个实际的道路分类问题。在这两种情况下,实验结果都证实了新方法几乎可以独立于初始条件产生最终解,并且分类性能显著提高。然后,我们使用多层网络,并将其应用于字符识别问题和政治数据分析。在这些问题中,我们可以证明通过将输入模式上的信息含量减少到某些点来提高噪声容忍性能。
In this paper, we propose a new information theoretic competitive learning method. We first construct a learning method in single-layered networks, and then we extend it to supervised multi-layered networks. Competitive unit outputs are computed by the inverse of Euclidean distance between input patterns and connection weights. As distance is smaller, competitive unit outputs are stronger. In realizing competition, neither the winner-take-all algorithm nor the lateral inhibition is used. Instead, the new method is based upon mutual information maximization between input patterns and competitive units. In maximizing mutual information, the entropy of competitive units is increased as much as possible. This means that all competitive units must equally be used in our framework. Thus, no under-utilized neurons or dead neurons are generated. When using multi-layered networks, we can improve noise-tolerance performance by unifying information maximization and minimization. We applied our method with single-layered networks to a simple artificial data problem and an actual road classification problem. In both cases, experimental results confirmed that the new method can produce the final solutions almost independently of initial conditions, and classification performance is significantly improved. Then, we used multi-layered networks, and applied them to a character recognition problem and a political data analysis. In these problem, we could show that noise-tolerance performance was improved by decreasing information content on input patterns to certain points.