You Shouldn't Trust Me: Learning Models Which Conceal Unfairness From Multiple Explanation Methods

You Shouldn't Trust Me: Learning Models Which Conceal Unfairness From Multiple Explanation Methods
复制标题

DOI:
10.17863/cam.48825
复制
发表时间:
2020-01
期刊:
--
影响因子:
--
通讯作者:
B. Dimanov;Umang Bhatt;M. Jamnik;Adrian Weller
B. Dimanov;Umang Bhatt;M. Jamnik;Adrian Weller
中科院分区:
其他
文献类型:
--
作者:
B. Dimanov;Umang Bhatt;M. Jamnik;Adrian Weller

文献摘要

被引文献

相似文献

算法系统的透明度是一个重要的研究领域,它被认为是最终用户和监管机构对机器学习模型建立适当信任的一种方式。一种流行的方法,LIME [23],甚至建议模型解释可以回答“我为什么要相信你?”这个问题。在这里,我们展示了一种简单的方法,用于修改预先训练的模型,以操纵许多流行的特征重要性解释方法的输出,而准确性几乎没有变化,从而证明了信任这种解释方法的危险。我们展示了这种解释攻击如何掩盖模型对敏感特征的歧视性使用,从而引起了人们对使用这种解释方法来检查模型公平性的强烈关注。
Transparency of algorithmic systems is an important area of research, which has been discussed as a way for end-users and regulators to develop appropriate trust in machine learning models. One popular approach, LIME [23], even suggests that model expla- nations can answer the question “Why should I trust you?”. Here we show a straightforward method for modifying a pre-trained model to manipulate the output of many popular feature importance explana- tion methods with little change in accuracy, thus demonstrating the danger of trusting such explanation methods. We show how this ex- planation attack can mask a model’s discriminatory use of a sensitive feature, raising strong concerns about using such explanation meth- ods to check fairness of a model.