Clustering units in neural networks: upstream vs downstream information

Clustering units in neural networks: upstream vs downstream information
复制标题

DOI:
10.48550/arxiv.2203.11815
复制
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Richard D. Lange;D. Rolnick;K. Kording
Richard D. Lange;D. Rolnick;K. Kording
中科院分区:
其他
文献类型:
--
作者:
Richard D. Lange;D. Rolnick;K. Kording

文献摘要

被引文献

相似文献

有人假设,人工神经网络中的某种形式的“模块化”结构应该对学习、组合和泛化有用。然而,定义和量化模块化仍然是一个悬而未决的问题。我们将检测功能模块的问题转化为检测功能相似单元的集群的问题。这就引出了一个问题,是什么让两个单元在功能上相似。为此,我们考虑两大类方法:那些定义相似性的基础上如何单位响应结构化的变化输入(“上游”),以及那些基于隐藏的单位激活的变化如何影响输出(“下游”)。我们进行了一项实证研究,量化模块化的隐藏层表示的简单前馈,全连接的网络,在一系列的超参数。对于每个模型,我们使用各种上游和下游措施来量化每层隐藏单元之间的成对关联,然后使用网络科学中的既定工具通过最大化其“模块化得分”来对其进行聚类。我们发现了两个令人惊讶的结果:第一,dropout极大地增加了模块化,而其他形式的权重正则化的效果更温和。其次,虽然我们观察到,通常有很好的协议内的上游方法和下游方法的集群,有一点协议的集群分配在这两个家庭的方法。这对表征学习具有重要意义,因为它表明,找到反映输入结构的模块化表征(例如解纠缠)可能是学习反映输出结构的模块化表征(例如组合性)的不同目标。
It has been hypothesized that some form of"modular"structure in artificial neural networks should be useful for learning, compositionality, and generalization. However, defining and quantifying modularity remains an open problem. We cast the problem of detecting functional modules into the problem of detecting clusters of similar-functioning units. This begs the question of what makes two units functionally similar. For this, we consider two broad families of methods: those that define similarity based on how units respond to structured variations in inputs ("upstream"), and those based on how variations in hidden unit activations affect outputs ("downstream"). We conduct an empirical study quantifying modularity of hidden layer representations of simple feedforward, fully connected networks, across a range of hyperparameters. For each model, we quantify pairwise associations between hidden units in each layer using a variety of both upstream and downstream measures, then cluster them by maximizing their"modularity score"using established tools from network science. We find two surprising results: first, dropout dramatically increased modularity, while other forms of weight regularization had more modest effects. Second, although we observe that there is usually good agreement about clusters within both upstream methods and downstream methods, there is little agreement about the cluster assignments across these two families of methods. This has important implications for representation-learning, as it suggests that finding modular representations that reflect structure in inputs (e.g. disentanglement) may be a distinct goal from learning modular representations that reflect structure in outputs (e.g. compositionality).