Towards Understanding Fairness and its Composition in Ensemble Machine Learning

Towards Understanding Fairness and its Composition in Ensemble Machine Learning
复制标题

DOI:
10.1109/icse48619.2023.00133
复制
发表时间:
2022-12
期刊:
2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Usman Gohar;Sumon Biswas;Hridesh Rajan
Usman Gohar;Sumon Biswas;Hridesh Rajan
中科院分区:
其他
文献类型:
--
作者:
Usman Gohar;Sumon Biswas;Hridesh Rajan

文献摘要

相似文献

机器学习(ML)软件在现代社会中得到了广泛的应用,据报道,机器学习对基于种族、性别、年龄等的少数群体具有公平性。最近的许多工作提出了测量和减轻ML模型中算法偏差的方法。现有的方法主要集中在基于单分类器的最大似然模型上。然而,现实世界的ML模型通常是由多个独立或依赖的学习者在一个集成(例如,随机森林)中组成的,其中公平性以一种非平凡的方式构成。公平是如何在整体上构成的?学习者的公平对合奏的最终公平有何影响?公平的学习者会导致不公平的集体吗?此外,研究表明,超参数会影响ML模型的公平性。集合超参数更复杂,因为它们影响学习者如何在不同类别的集合中组合。了解集合超参数对公平性的影响将有助于程序员设计公平的集合。今天,对于不同的集成算法,我们还不能完全理解这些。在本文中,我们全面研究了现实世界中流行的合奏:装袋、助推、堆叠和投票。我们在四个流行的公平数据集上建立了一个基准,包括从Kaggle收集的168个整体模型。我们使用现有的公平度量来理解公平的构成。我们的结果表明,在不使用缓解技术的情况下,合奏可以被设计得更公平。我们还确定了公平构成和数据特征之间的相互作用,以指导公平集成设计。最后,我们的基准可以被用来进一步研究公平的合奏。据我们所知,这是文献中第一次也是最大规模的关于合奏中公平组成的研究之一。
Machine Learning (ML) software has been widely adopted in modern society, with reported fairness implications for minority groups based on race, sex, age, etc. Many recent works have proposed methods to measure and mitigate algorithmic bias in ML models. The existing approaches focus on single classifier-based ML models. However, real-world ML models are often composed of multiple independent or dependent learners in an ensemble (e.g., Random Forest), where the fairness composes in a non-trivial way. How does fairness compose in ensembles? What are the fairness impacts of the learners on the ultimate fairness of the ensemble? Can fair learners result in an unfair ensemble? Furthermore, studies have shown that hyperparameters influence the fairness of ML models. Ensemble hyperparameters are more complex since they affect how learners are combined in different categories of ensembles. Understanding the impact of ensemble hyperparameters on fairness will help programmers design fair ensembles. Today, we do not understand these fully for different ensemble algorithms. In this paper, we comprehensively study popular real-world ensembles: Bagging, Boosting, Stacking, and Voting. We have developed a benchmark of 168 ensemble models collected from Kaggle on four popular fairness datasets. We use existing fairness metrics to understand the composition of fairness. Our results show that ensembles can be designed to be fairer without using mitigation techniques. We also identify the interplay between fairness composition and data characteristics to guide fair ensemble design. Finally, our benchmark can be leveraged for further research on fair ensembles. To the best of our knowledge, this is one of the first and largest studies on fairness composition in ensembles yet presented in the literature.