The impact of inconsistent human annotations on AI driven clinical decision making.

The impact of inconsistent human annotations on AI driven clinical decision making.
复制标题

DOI:
10.1038/s41746-023-00773-3
复制
发表时间:
2023-02-21
影响因子:
15.2
通讯作者:
--
中科院分区:
医学1区
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

在监督学习模型开发中,通常使用领域专家来提供类标签(注释)。由于固有的专家偏见、判断和失误等因素,即使是经验丰富的临床专家对同一现象(例如医学图像、诊断或预后状态)进行注释时,通常也会出现注释不一致的情况。虽然它们的存在相对众所周知,但当监督学习应用于此类“嘈杂”标记数据时,在现实世界中,这种不一致的影响在很大程度上并未得到充分研究。为了阐明这些问题,我们对三个真实的重症监护病房 (ICU) 数据集进行了广泛的实验和分析。具体来说,各个模型是根据通用数据集构建的,由 11 名格拉斯哥伊丽莎白女王大学医院 ICU 顾问独立注释,并通过内部验证比较模型性能估计(Fleiss 的 κ = 0.383,即公平一致)。此外,在 HiRID 外部数据集上对这 11 个分类器进行了广泛的外部验证(在静态和时间序列数据集上),发现模型的分类具有较低的成对一致性(平均 Cohen’s κ = 0.255,即最小一致性)。此外,他们在做出出院决定(Fleiss’ κ = 0.174)方面的分歧往往大于预测死亡率(Fleiss’ κ = 0.267)。鉴于这些不一致之处,我们进行了进一步的分析,以评估当前获得黄金标准模型和确定共识的最佳实践。结果表明:(a) 在急性临床环境中可能并不总是存在“超级专家”(使用内部和外部验证模型性能作为代理); (b) 标准共识寻求(例如多数投票)始终导致模型不理想。然而,进一步的分析表明,评估注释的可学习性并仅使用“可学习的”注释数据集来确定共识在大多数情况下可以实现最佳模型。
In supervised learning model development, domain experts are often used to provide the class labels (annotations). Annotation inconsistencies commonly occur when even highly experienced clinical experts annotate the same phenomenon (e.g., medical image, diagnostics, or prognostic status), due to inherent expert bias, judgments, and slips, among other factors. While their existence is relatively well-known, the implications of such inconsistencies are largely understudied in real-world settings, when supervised learning is applied on such ‘noisy’ labelled data. To shed light on these issues, we conducted extensive experiments and analyses on three real-world Intensive Care Unit (ICU) datasets. Specifically, individual models were built from a common dataset, annotated independently by 11 Glasgow Queen Elizabeth University Hospital ICU consultants, and model performance estimates were compared through internal validation (Fleiss’ κ = 0.383 i.e., fair agreement). Further, broad external validation (on both static and time series datasets) of these 11 classifiers was carried out on a HiRID external dataset, where the models’ classifications were found to have low pairwise agreements (average Cohen’s κ = 0.255 i.e., minimal agreement). Moreover, they tend to disagree more on making discharge decisions (Fleiss’ κ = 0.174) than predicting mortality (Fleiss’ κ = 0.267). Given these inconsistencies, further analyses were conducted to evaluate the current best practices in obtaining gold-standard models and determining consensus. The results suggest that: (a) there may not always be a “super expert” in acute clinical settings (using internal and external validation model performances as a proxy); and (b) standard consensus seeking (such as majority vote) consistently leads to suboptimal models. Further analysis, however, suggests that assessing annotation learnability and using only ‘learnable’ annotated datasets for determining consensus achieves optimal models in most cases.
DOI: 10.1016/j.media.2020.101759
发表时间: 2020-10-01
影响因子: 10.9
作者:
Karimi, Davood;Dou, Haoran;Gholipour, Ali
通讯作者: Gholipour, Ali
DOI: 10.1038/sdata.2016.35
发表时间: 2016-05-24
期刊: Scientific data
影响因子: 9.8
作者:
Johnson AE;Pollard TJ;Shen L;Lehman LW;Feng M;Ghassemi M;Moody B;Szolovits P;Celi LA;Mark RG
通讯作者: Mark RG
DOI: 10.7326/m14-0697
发表时间: 2015-02-01
影响因子: 7.2
作者:
Collins, Gary S.;Reitsma, Johannes B.;Moons, Karel G. M.
通讯作者: Moons, Karel G. M.
DOI: 10.1016/j.clinph.2014.11.008
发表时间: 2015-09-01
影响因子: 4.7
作者:
Halford, J. J.;Shiau, D.;LaRoche, S. M.
通讯作者: LaRoche, S. M.
DOI: 10.1186/1471-2288-14-40
发表时间: 2014-03-19
影响因子: 4
作者:
Collins GS;de Groot JA;Dutton S;Omar O;Shanyinde M;Tajar A;Voysey M;Wharton R;Yu LM;Moons KG;Altman DG
通讯作者: Altman DG