Fairwashing: the risk of rationalization

Fairwashing: the risk of rationalization
复制标题

DOI:
--
复制
发表时间:
2019-01
期刊:
--
影响因子:
--
通讯作者:
U. Aïvodji;Hiromi Arai;O. Fortineau;S. Gambs;Satoshi Hara;Alain Tapp
U. Aïvodji;Hiromi Arai;O. Fortineau;S. Gambs;Satoshi Hara;Alain Tapp
中科院分区:
其他
文献类型:
--
作者:
U. Aïvodji;Hiromi Arai;O. Fortineau;S. Gambs;Satoshi Hara;Alain Tapp

文献摘要

被引文献

相似文献

黑盒解释是解释机器学习模型如何产生结果的问题,其内部逻辑对审计人员来说是隐藏的,通常是复杂的。目前解决这一问题的方法有模型解释、结果解释和模型检验。虽然这些技术可以通过提供可解释性而受益,但它们可以以负面的方式用于执行公平清洗,我们将其定义为促进机器学习模型尊重某些道德价值观的错误看法。特别是,我们证明,它是可能的,系统合理化的决策,由一个不公平的黑箱模型使用模型的解释,以及结果的解释方法与给定的公平性度量。我们的解决方案,LaundryML,是基于正则化规则列表枚举算法,其目标是寻找公平的规则列表近似不公平的黑盒模型。我们在真实世界数据集上训练的黑盒模型上实证评估了我们的合理化技术,并表明可以获得对黑盒模型具有高保真度的规则列表,同时大大减少不公平。
Black-box explanation is the problem of explaining how a machine learning model -- whose internal logic is hidden to the auditor and generally complex -- produces its outcomes. Current approaches for solving this problem include model explanation, outcome explanation as well as model inspection. While these techniques can be beneficial by providing interpretability, they can be used in a negative manner to perform fairwashing, which we define as promoting the false perception that a machine learning model respects some ethical values. In particular, we demonstrate that it is possible to systematically rationalize decisions taken by an unfair black-box model using the model explanation as well as the outcome explanation approaches with a given fairness metric. Our solution, LaundryML, is based on a regularized rule list enumeration algorithm whose objective is to search for fair rule lists approximating an unfair black-box model. We empirically evaluate our rationalization technique on black-box models trained on real-world datasets and show that one can obtain rule lists with high fidelity to the black-box model while being considerably less unfair at the same time.