Auditing Black-Box Models for Indirect Influence

Auditing Black-Box Models for Indirect Influence
复制标题

审计黑盒模型的间接影响

DOI:
10.1109/icdm.2016.0011
复制
发表时间:
2016
期刊:
IEEE 16th International Conference on Data Mining (ICDM
影响因子:
--
通讯作者:
Venkatasubramanian, Suresh
Venkatasubramanian, Suresh
中科院分区:
--
文献类型:
--
作者:
Adler, Philip;Falk, Casey;Friedler, Sorelle A.;Rybeck, Gabriel;Scheidegger, Carlos;Smith, Brandon;Venkatasubramanian, Suresh

文献摘要

相似文献

数据训练的预测模型得到了广泛的使用,但在大多数情况下,它们被用作输出预测或分数的黑匣子。因此,很难更深入地了解模型行为,特别是不同的特征如何影响模型预测。这在解释复杂模型的行为或断言某些有问题的属性(如种族或性别)不会适时地影响决策时很重要。在本文中,我们提出了一种技术forcrowningblack-box模型,它可以让我们研究现有的模型在多大程度上利用数据集中的特定功能,而不知道模型是如何工作的。我们的工作重点是间接影响的问题:一些特征如何通过其他相关特征间接影响结果。因此,即使在进一步直接检查模型时,模型根本没有引用属性的情况下,我们也可以发现属性影响。我们的方法不需要重新训练黑盒模型。例如,如果模型只能通过API访问,这一点很重要,并且将我们的工作与其他研究特征影响的方法(如特征选择)进行了对比。我们提出了实验证据,我们的程序使用各种公开的数据集和模型的有效性。我们还使用可解释学习和特征选择技术以及其他黑盒审计程序来验证我们的程序。为了进一步证明这种技术的有效性,我们用它来审计一个黑箱累犯预测算法。
Data-trained predictive models see widespread use, but for the most part they are used asblack boxeswhich output a prediction or score. It is therefore hard to acquire a deeper understanding of model behavior and in particular how different features influence the model prediction. This is important when interpreting the behavior of complex models or asserting that certain problematic attributes (such as race or gender) arenotunduly influencing decisions. In this paper, we present a technique forauditingblack-box models, which lets us study the extent to which existing models take advantage of particular features in the data set, without knowing how the models work. Our work focuses on the problem ofindirect influence: how some features might indirectly influence outcomes via other, related features. As a result, we can find attribute influences even in cases where, upon further direct examination of the model,the attribute is not referred to by the model at all.Our approach does not require the black-box model to be retrained. This is important if, for example, the model is only accessible via an API, and contrasts our work with other methods that investigate feature influence such as feature selection. We present experimental evidence for the effectiveness of our procedure using a variety of publicly available data sets and models. We also validate our procedure using techniques from interpretable learning and feature selection, as well as against other black-box auditing procedures. To further demonstrate the effectiveness of this technique, we use it to audit a black-box recidivism prediction algorithm.