Towards Intersectionality in Machine Learning: Including More Identities, Handling Underrepresentation, and Performing Evaluation

Towards Intersectionality in Machine Learning: Including More Identities, Handling Underrepresentation, and Performing Evaluation
复制标题

迈向机器学习的交叉性:包括更多身份、处理代表性不足以及进行评估

DOI:
10.1145/3531146.3533101
复制
发表时间:
2022
期刊:
and Transparency
影响因子:
--
通讯作者:
Russakovsky, Olga
Russakovsky, Olga
中科院分区:
--
文献类型:
--
作者:
Wang, Angelina;Ramaswamy, Vikram V;Russakovsky, Olga

文献摘要

参考文献

被引文献

相似文献

机器学习公平性的研究历来被认为是一个单一的二元人口统计属性;然而,现实当然要复杂得多。在这项工作中,我们努力解决在机器学习管道的三个阶段出现的问题,当将交叉性作为多个人口统计属性合并时:(1)将哪些人口统计属性包括为数据集标签,(2)如何在模型训练期间处理逐渐缩小的子组规模,以及(3)如何超越现有的评估指标,为更多的子组对模型公平性进行基准测试。对于每个问题,我们对来自美国人口普查的表格数据集进行了彻底的实证评估,并为机器学习社区提出了建设性的建议。首先,我们主张在选择要训练的人口统计属性标签时,用经验验证来补充领域知识,同时始终对完整的人口统计属性集进行评估。其次,我们警告不要在不考虑其规范含义的情况下使用数据不平衡技术,并建议使用数据结构的替代方法。第三,我们引入了新的评估指标,更适合于交叉设置。总的来说,在将交叉性纳入机器学习时,我们就三个必要的(尽管还不够!)考虑因素提供了实质性的建议。
Research in machine learning fairness has historically considered a single binary demographic attribute; however, the reality is of course far more complicated. In this work, we grapple with questions that arise along three stages of the machine learning pipeline when incorporating intersectionality as multiple demographic attributes: (1) which demographic attributes to include as dataset labels, (2) how to handle the progressively smaller size of subgroups during model training, and (3) how to move beyond existing evaluation metrics when benchmarking model fairness for more subgroups. For each question, we provide thorough empirical evaluation on tabular datasets derived from the US Census, and present constructive recommendations for the machine learning community. First, we advocate for supplementing domain knowledge with empirical validation when choosing which demographic attribute labels to train on, while always evaluating on the full set of demographic attributes. Second, we warn against using data imbalance techniques without considering their normative implications and suggest an alternative using the structure in the data. Third, we introduce new evaluation metrics which are more appropriate for the intersectional setting. Overall, we provide substantive suggestions on three necessary (albeit not sufficient!) considerations when incorporating intersectionality into machine learning.
DOI: --
发表时间: 2018-11
期刊: arXiv: Applications
影响因子: --
作者:
Shira Mitchell;E. Potash;Solon Barocas
通讯作者: Shira Mitchell;E. Potash;Solon Barocas
计算机视觉的分析潜力和计算经验主义的挑战
DOI: 10.1145/3287560.3287568
发表时间: 2019
期刊: Proceedings of the 2019 ACM FAT* Conference
影响因子: --
作者:
Goldenfein, Jake
通讯作者: Goldenfein, Jake
DOI: 10.1109/cvpr46437.2021.00918
发表时间: 2020-12
期刊: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子: --
作者:
V. V. Ramaswamy-V.;Sunnie S. Y. Kim;Olga Russakovsky
通讯作者: V. V. Ramaswamy-V.;Sunnie S. Y. Kim;Olga Russakovsky
DOI: --
发表时间: 2018
期刊:
影响因子: --
作者:
Anna Carastathis
通讯作者: Anna Carastathis
最大最小准则的一些原因
DOI: 10.2307/j.ctv1cbn3j4.14
发表时间: 1974
期刊: The American Economic Review
影响因子: --
作者:
J. Rawls
通讯作者: J. Rawls