Grounding Representation Similarity with Statistical Testing

Grounding Representation Similarity with Statistical Testing
复制标题

DOI:
--
复制
发表时间:
2021-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Frances Ding;Jean-Stanislas Denain;J. Steinhardt
Frances Ding;Jean-Stanislas Denain;J. Steinhardt
中科院分区:
其他
文献类型:
--
作者:
Frances Ding;Jean-Stanislas Denain;J. Steinhardt

文献摘要

相似文献

为了理解神经网络行为,最近的工作使用规范相关分析(CCA)、中心核对齐(CKA)和其他相异性度量来定量比较不同网络的学习表示。不幸的是,这些广泛使用的措施通常在基本观察上存在分歧,例如仅在随机初始化方面不同的深度网络是否学习相似的表示。这些分歧提出了一个问题:我们应该相信这些差异性衡量标准中的哪一个(如果有的话)?我们提供了一个框架,通过具体的测试来解决这个问题:措施应该对影响功能行为的变化具有敏感性,并对不影响功能行为的变化具有特异性。我们通过各种功能行为来量化这一点,包括探测准确性和对分布变化的鲁棒性,并检查变化,例如改变随机初始化和删除主成分。我们发现当前指标表现出不同的弱点,注意到经典基线的表现出人意料地好,并突出显示所有指标似乎都失败的设置,从而为进一步改进提供了挑战集。
To understand neural network behavior, recent works quantitatively compare different networks' learned representations using canonical correlation analysis (CCA), centered kernel alignment (CKA), and other dissimilarity measures. Unfortunately, these widely used measures often disagree on fundamental observations, such as whether deep networks differing only in random initialization learn similar representations. These disagreements raise the question: which, if any, of these dissimilarity measures should we believe? We provide a framework to ground this question through a concrete test: measures should have sensitivity to changes that affect functional behavior, and specificity against changes that do not. We quantify this through a variety of functional behaviors including probing accuracy and robustness to distribution shift, and examine changes such as varying random initialization and deleting principal components. We find that current metrics exhibit different weaknesses, note that a classical baseline performs surprisingly well, and highlight settings where all metrics appear to fail, thus providing a challenge set for further improvement.