Combining Feature and Instance Attribution to Detect Artifacts

Combining Feature and Instance Attribution to Detect Artifacts
复制标题

DOI:
10.18653/v1/2022.findings-acl.153
复制
发表时间:
2021-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Pouya Pezeshkpour;Sarthak Jain;Sameer Singh;Byron C. Wallace
Pouya Pezeshkpour;Sarthak Jain;Sameer Singh;Byron C. Wallace
中科院分区:
其他
文献类型:
--
作者:
Pouya Pezeshkpour;Sarthak Jain;Sameer Singh;Byron C. Wallace

文献摘要

相似文献

训练主导 NLP 的深度神经网络需要大量数据集。这些通常是自动收集或通过众包收集的,并且可能表现出系统偏差或注释伪影。后者是指输入和输出之间的虚假相关性,并不代表特征和类别之间普遍存在的因果关系;利用这种相关性的模型可能看起来可以很好地执行给定的任务,但在样本数据外却会失败。在本文中,我们评估了使用不同的归因方法来帮助识别训练数据工件。我们提出了新的混合方法,将显着图(突出重要的输入特征)与实例归因方法(检索对给定预测有影响的训练样本)结合起来。我们证明,当存在具有挑战性的验证集时,所提出的训练特征归因可用于有效地发现训练数据中的伪影。我们还进行了一项小型用户研究,以评估这些方法在实践中是否对 NLP 研究人员有用,并取得了有希望的结果。我们为本文中的所有方法和实验提供了代码。
Training the deep neural networks that dominate NLP requires large datasets. These are often collected automatically or via crowdsourcing, and may exhibit systematic biases or annotation artifacts. By the latter we mean spurious correlations between inputs and outputs that do not represent a generally held causal relationship between features and classes; models that exploit such correlations may appear to perform a given task well, but fail on out of sample data. In this paper, we evaluate use of different attribution methods for aiding identification of training data artifacts. We propose new hybrid approaches that combine saliency maps (which highlight important input features) with instance attribution methods (which retrieve training samples influential to a given prediction). We show that this proposed training-feature attribution can be used to efficiently uncover artifacts in training data when a challenging validation set is available. We also carry out a small user study to evaluate whether these methods are useful to NLP researchers in practice, with promising results. We make code for all methods and experiments in this paper available.