Towards Debiasing DNN Models from Spurious Feature Influence

Towards Debiasing DNN Models from Spurious Feature Influence
复制标题

DOI:
10.1609/aaai.v36i9.21185
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Mengnan Du;Ruixiang Tang;Weijie Fu;Xia Hu
Mengnan Du;Ruixiang Tang;Weijie Fu;Xia Hu
中科院分区:
其他
文献类型:
--
作者:
Mengnan Du;Ruixiang Tang;Weijie Fu;Xia Hu

文献摘要

相似文献

最近的研究表明,深度神经网络(DNN)倾向于对某些人口群体表现出歧视。我们观察到,算法歧视可以解释的高度依赖模型的公平敏感的功能。出于这一观察的动机,我们建议通过抑制DNN模型捕获这些公平敏感特征与底层任务之间的虚假相关来实现公平性。具体来说,我们首先训练了一个只存在偏见的教师模型,该模型被明确鼓励最大限度地使用公平敏感特征进行预测。教师模型然后反教导去偏见的学生模型,使得学生模型的解释与教师模型的解释正交。关键思想是,由于教师模型明确依赖于公平敏感特征进行预测,正交解释损失迫使学生网络减少对敏感特征的依赖,而是捕获更多与任务相关的特征进行预测。实验分析表明,我们的框架大大减少了模型的公平敏感功能的关注。在四个数据集上的实验结果进一步验证了我们的框架在三个组公平性度量方面始终提高了公平性,具有相当甚至更好的准确性。
Recent studies indicate that deep neural networks (DNNs) are prone to show discrimination towards certain demographic groups. We observe that algorithmic discrimination can be explained by the high reliance of the models on fairness sensitive features. Motivated by this observation, we propose to achieve fairness by suppressing the DNN models from capturing the spurious correlation between those fairness sensitive features with the underlying task. Specifically, we firstly train a bias-only teacher model which is explicitly encouraged to maximally employ fairness sensitive features for prediction. The teacher model then counter-teaches a debiased student model so that the interpretation of the student model is orthogonal to the interpretation of the teacher model. The key idea is that since the teacher model relies explicitly on fairness sensitive features for prediction, the orthogonal interpretation loss enforces the student network to reduce its reliance on sensitive features and instead capture more task relevant features for prediction. Experimental analysis indicates that our framework substantially reduces the model's attention on fairness sensitive features. Experimental results on four datasets further validate that our framework has consistently improved the fairness with respect to three group fairness metrics, with a comparable or even better accuracy.