Addressing bias in prediction models by improving subpopulation calibration

Addressing bias in prediction models by improving subpopulation calibration
复制标题

DOI:
10.1093/jamia/ocaa283
复制
发表时间:
2021-03-01
影响因子:
6.4
通讯作者:
Dagan, Noa
Dagan, Noa
中科院分区:
管理学2区
文献类型:
--
作者:
Barda, Noam;Yona, Gal;Dagan, Noa

文献摘要

被引文献

相似文献

材料和方法:在这项回溯性队列研究中,我们评估了基于汇集队列方程(PCE)和骨折风险评估工具(FRAX)的预测在总体人群和由年龄、性别、种族、社会经济地位和移民史交叉定义的亚群中的校正情况。接下来,我们应用重新校准算法并评估校准指标的变化,包括整体校准。结果:PCE人群中有1021 041名患者,FRAX人群中有1116 324名患者。两个测试模型的基线总体模型校正良好,但在相当大一部分亚群中的校正较差。应用该算法后,子总体校正统计量有了很大的改善,PCE和FRAX模型的总体校正方差分别降低了98.8%和94.3%。校准,即预测的风险和观察到的风险之间的一致性,对于在模型开发集中代表性不足的子群体来说通常很差,导致这些子群体的偏见和业绩下降。在这项工作中,我们经验性地评估了Hebert-Johnson等人设计的公平算法的改进版本。结论:一种后处理和模型无关的预测模型重校准公平算法大大减少了子总体误校的偏差,从而提高了公平性和公平性。
Objective: To illustrate the problem of subpopulation miscalibration, to adapt an algorithm for recalibration of the predictions, and to validate its performance.Materials and Methods: In this retrospective cohort study, we evaluated the calibration of predictions based on the Pooled Cohort Equations (PCE) and the fracture risk assessment tool (FRAX) in the overall population and in subpopulations defined by the intersection of age, sex, ethnicity, socioeconomic status, and immigration history. We next applied the recalibration algorithm and assessed the change in calibration metrics, including calibration-in-the-large.Results: 1 021 041 patients were included in the PCE population, and 1 116 324 patients were included in the FRAX population. Baseline overall model calibration of the 2 tested models was good, but calibration in a substantial portion of the subpopulations was poor. After applying the algorithm, subpopulation calibration statistics were greatly improved, with the variance of the calibration-in-the-large values across all subpopulations reduced by 98.8% and 94.3% in the PCE and FRAX models, respectively.Discussion: Prediction models in medicine are increasingly common. Calibration, the agreement between predicted and observed risks, is commonly poor for subpopulations that were underrepresented in the development set of the models, resulting in bias and reduced performance for these subpopulations. In this work, we empirically evaluated an adapted version of the fairness algorithm designed by Hebert-Johnson et al. (2017) and demonstrated its use in improving subpopulation miscalibration.Conclusion: A postprocessing and model-independent fairness algorithm for recalibration of predictive models greatly decreases the bias of subpopulation miscalibration and thus increases fairness and equality.