Characterizing the risk of fairwashing

Characterizing the risk of fairwashing
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
U. Aïvodji;Hiromi Arai;S. Gambs;Satoshi Hara
U. Aïvodji;Hiromi Arai;S. Gambs;Satoshi Hara
中科院分区:
其他
文献类型:
--
作者:
U. Aïvodji;Hiromi Arai;S. Gambs;Satoshi Hara

文献摘要

相似文献

公平洗牌指的是,通过事后的解释操纵,不公平的黑箱模型可能会被更公平的模型解释。在本文中,我们通过分析它们的保真度-不公平性权衡来研究公平清洗攻击的能力。特别是,我们证明了公平的解释模型可以推广到起诉组之外(即被解释的数据点),这意味着公平的解释可以被用来合理化黑盒模型随后的不公平决定。我们还证明了公平清洗攻击可以在黑箱模型之间转移,这意味着其他黑箱模型可以在不显式使用他们的预测的情况下执行公平清洗。洗牌攻击的这种概括性和可转移性意味着在实践中很难检测到它们。最后,我们提出了一种基于高保真解说员不公平范围的量化洗牌风险的方法。
Fairwashing refers to the risk that an unfair black-box model can be explained by a fairer model through post-hoc explanation manipulation. In this paper, we investigate the capability of fairwashing attacks by analyzing their fidelity-unfairness trade-offs. In particular, we show that fairwashed explanation models can generalize beyond the suing group (i.e., data points that are being explained), meaning that a fairwashed explainer can be used to rationalize subsequent unfair decisions of a black-box model. We also demonstrate that fairwashing attacks can transfer across black-box models, meaning that other black-box models can perform fairwashing without explicitly using their predictions. This generalization and transferability of fairwashing attacks imply that their detection will be difficult in practice. Finally, we propose an approach to quantify the risk of fairwashing, which is based on the computation of the range of the unfairness of high-fidelity explainers.