Automated Directed Fairness Testing

Automated Directed Fairness Testing
复制标题

DOI:
10.1145/3238147.3238165
复制
发表时间:
2018-07
期刊:
2018 33rd IEEE/ACM International Conference on Automated Software Engineering (ASE)
影响因子:
--
通讯作者:
Sakshi Udeshi;Pryanshu Arora;Sudipta Chattopadhyay
Sakshi Udeshi;Pryanshu Arora;Sudipta Chattopadhyay
中科院分区:
其他
文献类型:
--
作者:
Sakshi Udeshi;Pryanshu Arora;Sudipta Chattopadhyay

文献摘要

被引文献

相似文献

公正性是决策过程中的一个重要特征。随着机器学习模型越来越多地被用于敏感的应用领域(如教育和就业)进行决策,由这种模型计算的决策必须没有意外的偏见。但是,我们如何自动验证任意机器学习模型的公平性呢?对于给定的机器学习模型和一组敏感的输入参数,我们的Aequitas方法会自动发现突出违反公平的歧视性输入。Aequitas的核心是三种新的策略,即在输入空间上使用概率搜索,目的是揭露违反公平的行为。我们的Aequitas方法利用常见机器学习模型中固有的健壮性属性来设计和实现可伸缩的测试生成方法。我们生成的测试输入的一个吸引人的特征是,它们可以系统地添加到底层模型的训练集,并提高其公平性。为此,我们设计了一个全自动化的模块,保证了模型的公平性。我们实现了Aequitas,并在六个最先进的分类器上对其进行了评估。我们的主题还包括一个在设计时考虑到公平的分类器。我们表明,Aequitas有效地生成了输入来发现所有主题分类器中的公平性违规行为,并使用生成的测试输入系统地提高了各个模型的公平性。在我们的评估中,Aequitas产生了高达70%的歧视性投入(w.r.t.生成的输入总数),并利用这些输入将公平性提高高达94%。
Fairness is a critical trait in decision making. As machine-learning models are increasingly being used in sensitive application domains (e.g. education and employment) for decision making, it is crucial that the decisions computed by such models are free of unintended bias. But how can we automatically validate the fairness of arbitrary machine-learning models? For a given machine-learning model and a set of sensitive input parameters, our Aequitas approach automatically discovers discriminatory inputs that highlight fairness violation. At the core of Aequitas are three novel strategies to employ probabilistic search over the input space with the objective of uncovering fairness violation. Our Aequitas approach leverages inherent robustness property in common machine-learning models to design and implement scalable test generation methodologies. An appealing feature of our generated test inputs is that they can be systematically added to the training set of the underlying model and improve its fairness. To this end, we design a fully automated module that guarantees to improve the fairness of the model. We implemented Aequitas and we have evaluated it on six state-of-the-art classifiers. Our subjects also include a classifier that was designed with fairness in mind. We show that Aequitas effectively generates inputs to uncover fairness violation in all the subject classifiers and systematically improves the fairness of respective models using the generated test inputs. In our evaluation, Aequitas generates up to 70% discriminatory inputs (w.r.t. the total number of inputs generated) and leverages these inputs to improve the fairness up to 94%.