Accuracy comparison across face recognition algorithms: Where are we on measuring race bias?

Accuracy comparison across face recognition algorithms: Where are we on measuring race bias?
复制标题

DOI:
10.1109/tbiom.2020.3027269
复制
发表时间:
2020-09-29
期刊:
IEEE transactions on biometrics, behavior, and identity science
影响因子:
--
通讯作者:
O’Toole AJ
O’Toole AJ
中科院分区:
其他
文献类型:
--
作者:
Cavazos JG;Phillips PJ;Castillo CD;O’Toole AJ

文献摘要

被引文献

相似文献

前几代人脸识别算法对不同种族图像的准确性不同(种族偏见)。在这里,我们提出了可能的潜在因素(数据驱动和场景建模)和评估算法种族偏见的方法考虑。我们讨论数据驱动因素(例如,图像质量、图像总体统计和算法体系结构),以及考虑算法的“用户”角色的场景建模因素(例如,阈值决定和人口限制)。为了说明这些问题是如何应用的,我们展示了来自东亚和高加索面孔的四种人脸识别算法(上一代算法和三种深度卷积神经网络(DCNN))的数据。首先,数据集难度影响整体识别准确性和种族偏见,种族偏见随着项目难度的增加而增加。其次,对于所有四种算法,偏差的程度取决于识别决策阈值。为了实现相等的错误接受率(法尔斯),东亚面孔需要比高加索面孔更高的识别阈值,对于所有算法。第三,在测试中使用的分布公式的人口限制,影响算法准确性的估计。我们的结论是,种族偏见需要衡量个人的应用程序,我们提供了一个清单来衡量这种偏见的人脸识别算法。
Previous generations of face recognition algorithms differ in accuracy for images of different races (race bias). Here, we present the possible underlying factors (data-driven and scenario modeling) and methodological considerations for assessing race bias in algorithms. We discuss data-driven factors (e.g., image quality, image population statistics, and algorithm architecture), and scenario modeling factors that consider the role of the “user” of the algorithm (e.g., threshold decisions and demographic constraints). To illustrate how these issues apply, we present data from four face recognition algorithms (a previous-generation algorithm and three deep convolutional neural networks, DCNNs) for East Asian and Caucasian faces. First, dataset difficulty affected both overall recognition accuracy and race bias, such that race bias increased with item difficulty. Second, for all four algorithms, the degree of bias varied depending on the identification decision threshold. To achieve equal false accept rates (FARs), East Asian faces required higher identification thresholds than Caucasian faces, for all algorithms. Third, demographic constraints on the formulation of the distributions used in the test, impacted estimates of algorithm accuracy. We conclude that race bias needs to be measured for individual applications and we provide a checklist for measuring this bias in face recognition algorithms.