An Empirical Comparison of Instance Attribution Methods for NLP

An Empirical Comparison of Instance Attribution Methods for NLP
复制标题

DOI:
10.18653/v1/2021.naacl-main.75
复制
发表时间:
2021-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Pouya Pezeshkpour;Sarthak Jain;Byron C. Wallace;Sameer Singh
Pouya Pezeshkpour;Sarthak Jain;Byron C. Wallace;Sameer Singh
中科院分区:
其他
文献类型:
--
作者:
Pouya Pezeshkpour;Sarthak Jain;Byron C. Wallace;Sameer Singh

文献摘要

被引文献

相似文献

深度模型的广泛采用激发了对解释网络输出和促进模型调试的方法的迫切需求。实例归因方法是通过检索(可能)导致特定预测的训练实例来实现这些目标的一种方法。影响函数(IF; Koh和Liang 2017)通过量化扰动单个列车实例对特定测试预测的影响来提供实现这一点的机制。然而,即使近似IF也是计算上昂贵的,在许多情况下可能是禁止的。更简单的方法(例如,检索与给定测试点最相似的训练示例)执行训练?在这项工作中,我们评估的程度不同的潜在实例属性同意训练样本的重要性。我们发现,简单的检索方法产生的训练实例与通过基于梯度的方法(如IF)识别的训练实例不同,但仍然表现出与更复杂的归因方法相似的理想特征。本文中所有方法和实验的代码可在https://github.com/successar/instance_attributions_NLP上获得。
Widespread adoption of deep models has motivated a pressing need for approaches to interpret network outputs and to facilitate model debugging. Instance attribution methods constitute one means of accomplishing these goals by retrieving training instances that (may have) led to a particular prediction. Influence functions (IF; Koh and Liang 2017) provide machinery for doing this by quantifying the effect that perturbing individual train instances would have on a specific test prediction. However, even approximating the IF is computationally expensive, to the degree that may be prohibitive in many cases. Might simpler approaches (e.g., retrieving train examples most similar to a given test point) perform comparably? In this work, we evaluate the degree to which different potential instance attribution agree with respect to the importance of training samples. We find that simple retrieval methods yield training instances that differ from those identified via gradient-based methods (such as IFs), but that nonetheless exhibit desirable characteristics similar to more complex attribution methods. Code for all methods and experiments in this paper is available at: https://github.com/successar/instance_attributions_NLP.