Leveraging an Alignment Set in Tackling Instance-Dependent Label Noise

Leveraging an Alignment Set in Tackling Instance-Dependent Label Noise
复制标题

DOI:
10.48550/arxiv.2307.04868
复制
发表时间:
2023-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Donna Tjandra;J. Wiens
Donna Tjandra;J. Wiens
中科院分区:
其他
文献类型:
--
作者:
Donna Tjandra;J. Wiens

文献摘要

相似文献

嘈杂的训练标签会损害模型性能。大多数旨在解决标签噪声的方法都假设标签噪声与输入特征无关。然而,在实践中,标签噪声通常是特征或\textit{instance-dependent},因此是有偏差的(即,某些实例比其它实例更可能被错误标记)。例如,在一个示例中,在临床护理中,与男性患者相比,女性患者更有可能被诊断为心血管疾病。忽略这种依赖性的方法可能会产生具有较差区分性能的模型,并且在许多医疗保健环境中,可能会加剧健康差异的问题。鉴于这些限制,我们提出了一个两阶段的方法来学习存在实例相关的标签噪声。我们的方法利用\textit{\锚点},一个小的数据子集,我们知道观察到的和地面真实标签。在几项任务中,我们的方法在减少偏差(均衡优势曲线下的面积,AUEOC)的同时,使最先进的区分性能(AUROC)得到了持续的改善。例如,在MIMIC-III数据集上预测急性呼吸衰竭发作时,我们的方法实现了0.84(SD [标准差] 0.01)的调和平均值(AUROC和AUEOC),而下一个最佳基线的调和平均值为0.81(SD 0.01)。总的来说,我们的方法提高了准确性,同时减轻了潜在的偏见相比,现有的方法在存在实例相关的标签噪声。
Noisy training labels can hurt model performance. Most approaches that aim to address label noise assume label noise is independent from the input features. In practice, however, label noise is often feature or \textit{instance-dependent}, and therefore biased (i.e., some instances are more likely to be mislabeled than others). E.g., in clinical care, female patients are more likely to be under-diagnosed for cardiovascular disease compared to male patients. Approaches that ignore this dependence can produce models with poor discriminative performance, and in many healthcare settings, can exacerbate issues around health disparities. In light of these limitations, we propose a two-stage approach to learn in the presence instance-dependent label noise. Our approach utilizes \textit{\anchor points}, a small subset of data for which we know the observed and ground truth labels. On several tasks, our approach leads to consistent improvements over the state-of-the-art in discriminative performance (AUROC) while mitigating bias (area under the equalized odds curve, AUEOC). For example, when predicting acute respiratory failure onset on the MIMIC-III dataset, our approach achieves a harmonic mean (AUROC and AUEOC) of 0.84 (SD [standard deviation] 0.01) while that of the next best baseline is 0.81 (SD 0.01). Overall, our approach improves accuracy while mitigating potential bias compared to existing approaches in the presence of instance-dependent label noise.