Understanding and Mitigating Accuracy Disparity in Regression

Understanding and Mitigating Accuracy Disparity in Regression
复制标题

DOI:
--
复制
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Jianfeng Chi;Yuan Tian;Geoffrey J. Gordon;Han Zhao
Jianfeng Chi;Yuan Tian;Geoffrey J. Gordon;Han Zhao
中科院分区:
其他
文献类型:
--
作者:
Jianfeng Chi;Yuan Tian;Geoffrey J. Gordon;Han Zhao

文献摘要

相似文献

随着大规模预测系统在高风险领域的广泛应用,如人脸识别、刑事司法等,不同人口亚群之间预测精度的差异要求从根本上了解这种差异的来源,并通过算法干预来缓解这种差异。本文研究了回归中的精度差异问题。首先,我们提出了一个误差分解定理,将精度差异分解为边缘标签分布之间的距离和条件表示之间的距离,以帮助解释实际中出现这种精度差异的原因。基于这种误差分解和分布与统计距离对齐的一般思想,我们提出了一种减小这种差异的算法,并分析了它的博弈论最优目标函数。为了验证我们的理论发现,我们还在五个基准数据集上进行了实验。实验结果表明,本文提出的算法在保持回归模型预测能力的同时,有效地缓解了精度差距。
With the widespread deployment of large-scale prediction systems in high-stakes domains, e.g., face recognition, criminal justice, etc., disparity in prediction accuracy between different demographic subgroups has called for fundamental understanding on the source of such disparity and algorithmic intervention to mitigate it. In this paper, we study the accuracy disparity problem in regression. To begin with, we first propose an error decomposition theorem, which decomposes the accuracy disparity into the distance between marginal label distributions and the distance between conditional representations, to help explain why such accuracy disparity appears in practice. Motivated by this error decomposition and the general idea of distribution alignment with statistical distances, we then propose an algorithm to reduce this disparity, and analyze its game-theoretic optima of the proposed objective functions. To corroborate our theoretical findings, we also conduct experiments on five benchmark datasets. The experimental results suggest that our proposed algorithms can effectively mitigate accuracy disparity while maintaining the predictive power of the regression models.