Contextually guided unsupervised learning using local multivariate binary processors

Contextually guided unsupervised learning using local multivariate binary processors
复制标题

DOI:
10.1016/s0893-6080(97)00110-x
复制
发表时间:
1998-01-01
期刊:
影响因子:
7.8
通讯作者:
Phillips, WA
Phillips, WA
中科院分区:
计算机科学1区
文献类型:
--
作者:
Kay, J;Floreano, D;Phillips, WA

文献摘要

被引文献

相似文献

我们考虑上下文指导在多流神经网络学习和处理中的作用。早期的工作(Kay 和 Phillips,1994、1996;Phillips 等人,1995)展示了如何将特征发现和关联学习的目标融合在单个目标中,并使用信息论进行精确化,使得本地二进制处理器可以提取跨流一致的单个特征。在本文中,我们考虑具有多元二进制输出的多单元本地处理器,可以提取更多数量的相干特征。使用伊辛模型,我们定义了一类信息论目标函数和局部近似,并推导出两种情况下的学习规则。这些规则与著名的 BCM 规则有相似之处,也有不同之处。 infomax 的本地和全局版本以及相干 infomax 的多元版本都是通用方法的副产品。着眼于生物学上更合理的局部规则,我们描述了一些旨在研究处理器的特定属性和一般方法的计算实验。主要结论是:(1)本文介绍的本地方法具有所需的功能。 (2) 多单元处理器中的不同单元学会了对其感受野的不同方面做出反应。 (3) 每个处理器内的单元通常产生一个分布式代码,其中输出是相关的并且对损坏具有鲁棒性;在可用单元数量仅足以传递相关信息的特殊情况下,就产生了一种竞争性学习的形式。 (4) 上下文连接使得能够提取跨流相关的信息,并且通过改进弱或噪声输入的特征检测,它们在短期处理和提高泛化方面发挥了有用的作用。 (5)该方法允许学习分布式自组织群体代码之间的统计关联。 (C) 1998 Elsevier Science Ltd. 保留所有权利。
We consider the role of contextual guidance in learning and processing within multi-stream neural networks. Earlier work (Kay and Phillips, 1994, 1996; Phillips et al., 1995) showed how the goals of feature discovery and associative learning could be fused within a single objective and made precise using information theory in such a way that local binary processors could extract a single feature that is coherent across streams. In this paper, we consider multi-unit local processors with multivariate binary outputs that enable a greater number of coherent features to be extracted. Using the Ising model, we define a class of information-theoretic objective functions and also local approximations and derive the learning rules in both cases. These rules have similarities to, and differences from, the celebrated BCM rule. Local and global versions of infomax appear as by-products of the general approach, as well as multivariate versions of coherent infomax. Focussing on the more biologically plausible local rules, we describe some computational experiments designed to investigate specific properties of the processors and the general approach. The main conclusions are: (1) the local methodology introduced in the paper has the required functionality. (2) Different units within the multi-unit processors learned to respond to different aspects of their receptive fields. (3) The units within each processor generally produced a distributed code in which the outputs were correlated and which was robust to damage; in the special case where the number of units available was only just sufficient to transmit the relevant information, a form of competitive learning was produced. (4) The contextual connections enabled the information correlated across streams to be extracted and, by improving feature detection with weak or noisy inputs, they played a useful role in short-term processing and in improving generalization. (5) The methodology allows the statistical associations between distributed self-organizing population codes to be learned. (C) 1998 Elsevier Science Ltd. All rights reserved.