Equity in essence: a call for operationalising fairness in machine learning for healthcare.

Equity in essence: a call for operationalising fairness in machine learning for healthcare.
复制标题

DOI:
10.1136/bmjhci-2020-100289
复制
发表时间:
2021-04
影响因子:
4.1
通讯作者:
Ghassemi M
Ghassemi M
中科院分区:
其他
文献类型:
--
作者:
Wawira Gichoya J;McCoy LG;Celi LA;Ghassemi M

文献摘要

参考文献

被引文献

相似文献

用于医疗保健的机器学习(MLHC)正处于从期刊和会议记录的页面跳跃到床边临床实施的关键时刻。这一努力的成功需要综合来自机器学习和医疗保健领域的见解,以确保利用MLHC的独特特征来最大限度地提高效益并最大限度地降低风险。这一努力的一个重要部分是建立和正式确定这些工具的特点和评估其性能的过程和程序。在这方面取得的有意义的进展可以在最近制定的MLHC模型开发指南,1 MLHC临床试验设计和报告指南,2 3和MLHC工具的监管评估协议中找到。4 5但是,虽然这些准则和协议广泛涉及相关的技术考虑,但缺乏对公平性、偏见和意外的不同影响问题的处理。这些问题在更广泛的ML社区中占据了突出的位置,6-9最近的工作突出了面部识别和性别分类软件准确性的种族差异,6 - 10自然语言处理模型输出中的性别偏见,11 - 12以及保释和刑事判决算法中的种族偏见等问题。13 MLHC也不能免受这些问题的影响,如分配医疗资源的算法的不同结果所示,14 15在临床笔记16上开发的语言模型中存在偏见,以及主要在浅色皮肤图像上开发的黑色素瘤检测模型。[17]在本文中,我们将研究最近MLHC模型报告、临床试验和监管批准指南中的公平性。我们强调机会,以确保公平是根本的MLHC,并研究如何可以操作的MLHC背景下。公平作为一个事后的想法?模型开发和试验报告指南最近的几份文件试图列举MLHC的指导原则,这些原则具有不同程度的实际意义。总的来说,这些文档在突出人工智能(AI)特定的技术和操作问题方面做得很好,例如如何处理人与AI的交互,或者如何解释模型性能错误。然而,如表1所示,公平性的提法要么明显缺席,要么只是顺便提及,要么被降级为补充讨论。值得注意的例子是最近的标准方案项目:介入试验-AI(SPIRIT-AI)2和报告试验-AI(CONSORT-AI)3扩展的统一标准,这些扩展扩展了AI临床试验设计和报告的突出指南,以包括与AI相关的问题。虽然后者在讨论中指出,“还应鼓励研究者探索不同人群亚组之间的性能和错误率差异”,但指南本身并未正式纳入这一概念。类似地,即将发布的个体预后或诊断多变量预测模型的透明报告-ML(TRIPOD-ML)18和诊断准确性研究AI扩展报告标准(STARD-AI)19模型报告指南的公告文件也没有提到这些问题(尽管我们期待它们可能被纳入这些指南的最终版本)。虽然最近出版的指导方针,
INTRODUCTION Machine learning for healthcare (MLHC) is at the juncture of leaping from the pages of journals and conference proceedings to clinical implementation at the bedside. Succeeding in this endeavour requires the synthesis of insights from both the machine learning and healthcare domains, in order to ensure that the unique characteristics of MLHC are leveraged to maximise benefits and minimise risks. An important part of this effort is establishing and formalising processes and procedures for characterising these tools and assessing their performance. Meaningful progress in this direction can be found in recently developed guidelines for the development of MLHC models, 1 guidelines for the design and reporting of MLHC clinical trials, 2 3 and protocols for the regulatory assessment of MLHC tools. 4 5 But while such guidelines and protocols engage extensively with relevant technical considerations, engagement with issues of fairness, bias and unintended disparate impact is lacking. Such issues have taken on a place of prominence in the broader ML community, 6–9 with recent work highlighting issues such as racial disparities in the accuracy of facial recognition and gender classification software, 6 10 gender bias in the output of natural language processing models 11 12 and racial bias in algorithms for bail and criminal sentencing. 13 MLHC is not immune to these concerns, as seen in disparate outcomes from algorithms for allocating healthcare resources, 14 15 bias in language models developed on clinical notes 16 and melanoma detection models developed primarily on images of light-coloured skin. 17 Within this paper, we will examine the inclusion of fairness in recent guidelines for MLHC model reporting, clinical trials and regulatory approval. We highlight opportunities to ensure that fairness is made fundamental to MLHC, and examine ways how this can be operationalised for the MLHC context.FAIRNESS AS AN AFTERTHOUGHT? Model development and trial reporting guidelines Several recent documents have attempted, with varying degrees of practical implication, to enumerate guiding principles for MLHC. Broadly, these documents do an excellent job of highlighting artificial intelligence (AI)-specific technical and operational concerns, such as how to handle human-AI interaction, or how to account for model performance errors. Yet as outlined in table 1, references to fairness are either conspicuously absent, made merely in passing, or relegated to supplemental discussion. Notable examples are the recent the Standard Protocol Items: Recommendations for Interventional Trials-AI (SPIRIT-AI) 2 and Consolidated Standards of Reporting Trials-AI (CONSORT-AI) 3 extensions, which expand prominent guidelines for the design and reporting of AI clinical trials to include concerns relevant to AI. While the latter states in the discussion that ‘investigators should also be encouraged to explore differences in performance and error rates across population subgroups’, 3 there is no more formal inclusion of the concept into the guideline itself. Similarly, the announcement papers for the upcoming Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis-ML (TRIPOD-ML) 18 andStandards for Reporting of Diagnostic Accuracy Studies AI Extension (STARD-AI) 19 guidelines for model reporting do not allude to these issues (though we wait in anticipation for their potential inclusion in the final versions of these guidelines). While recently published guidelines from the editors of
DOI: 10.1093/jamia/ocaa133
发表时间: 2020-12-01
影响因子: 6.4
作者:
Ferryman, Kadija
通讯作者: Ferryman, Kadija
DOI: 10.1126/science.aal4230
发表时间: 2017-04-14
期刊: SCIENCE
影响因子: 56.9
作者:
Caliskan, Aylin;Bryson, Joanna J.;Narayanan, Arvind
通讯作者: Narayanan, Arvind
DOI: 10.1126/science.aax2342
发表时间: 2019-10-25
期刊: SCIENCE
影响因子: 56.9
作者:
Obermeyer, Ziad;Powers, Brian;Mullainathan, Sendhil
通讯作者: Mullainathan, Sendhil
DOI: 10.1097/ccm.0000000000004246
发表时间: 2020-05-01
影响因子: 8.8
作者:
Leisman, Daniel E.;Harhay, Michael O.;Maslove, David M.
通讯作者: Maslove, David M.
DOI: 10.1038/s41591-020-1037-7
发表时间: 2020-09
期刊: Nature medicine
影响因子: 82.9
作者:
Cruz Rivera S;Liu X;Chan AW;Denniston AK;Calvert MJ;SPIRIT-AI and CONSORT-AI Working Group;SPIRIT-AI and CONSORT-AI Steering Group;SPIRIT-AI and CONSORT-AI Consensus Group
通讯作者: SPIRIT-AI and CONSORT-AI Consensus Group