Accounting for Label Uncertainty in Machine Learning for Detection of Acute Respiratory Distress Syndrome.

Accounting for Label Uncertainty in Machine Learning for Detection of Acute Respiratory Distress Syndrome.
复制标题

DOI:
10.1109/jbhi.2018.2810820
复制
发表时间:
2019-01
影响因子:
7.7
通讯作者:
Najarian K
Najarian K
中科院分区:
工程技术1区
文献类型:
--
作者:
Reamaroon N;Sjoding MW;Lin K;Iwashyna TJ;Najarian K

文献摘要

被引文献

相似文献

当在一些临床应用中训练用于监督学习任务的机器学习算法时,一些患者的正确标签的不确定性可能会对算法的性能产生不利影响。例如,即使是临床专家在为一些患者分配医学诊断时也可能由于患者病例的模糊性或诊断标准的不完全可靠性而具有较低的置信度。因此,在算法训练中使用的一些情况可能被错误标记,从而对算法的性能产生不利影响。然而,专家也可以量化他们在这些情况下的诊断不确定性。我们提出了一个强大的方法实现支持向量机,占这样的临床诊断的不确定性时,训练的算法来检测患者发展的急性呼吸窘迫综合征(ARDS)。急性呼吸窘迫综合征是一种危重病人的综合征,其诊断使用已知不完善的临床标准。我们将ARDS诊断的不确定性表示为与每个训练标签相关的分级置信度权重。我们还进行了一种新的时间序列采样方法,以解决模型训练中使用的每个患者的纵向临床数据之间的相互关联问题,以限制过拟合。初步结果表明,我们可以实现有意义的改进算法的性能,以检测患者与ARDS的一个hold-out样本,当我们比较我们的方法,占训练标签的不确定性与传统的SVM算法。
When training a machine learning algorithm for a supervised-learning task in some clinical applications, uncertainty in the correct labels of some patients may adversely affect the performance of the algorithm. For example, even clinical experts may have less confidence when assigning a medical diagnosis to some patients because of ambiguity in the patient’s case or imperfect reliability of the diagnostic criteria. As a result, some cases used in algorithm training may be mis-labeled, adversely affecting the algorithm’s performance. However, experts may also be able to quantify their diagnostic uncertainty in these cases. We present a robust method implemented with Support Vector Machines to account for such clinical diagnostic uncertainty when training an algorithm to detect patients who develop the acute respiratory distress syndrome (ARDS). ARDS is a syndrome of the critically ill that is diagnosed using clinical criteria known to be imperfect. We represent uncertainty in the diagnosis of ARDS as a graded weight of confidence associated with each training label. We also performed a novel time-series sampling method to address the problem of inter-correlation among the longitudinal clinical data from each patient used in model training to limit overfitting. Preliminary results show that we can achieve meaningful improvement in the performance of algorithm to detect patients with ARDS on a hold-out sample, when we compare our method that accounts for the uncertainty of training labels with a conventional SVM algorithm.