Reducing False-Positive Results in Newborn Screening Using Machine Learning

Reducing False-Positive Results in Newborn Screening Using Machine Learning
复制标题

DOI:
10.3390/ijns6010016
复制
发表时间:
2020-03-01
影响因子:
3.5
通讯作者:
Scharfe, Curt
Scharfe, Curt
中科院分区:
其他
文献类型:
--
作者:
Peng, Gang;Tang, Yishuo;Scharfe, Curt

文献摘要

被引文献

相似文献

新生儿先天代谢紊乱筛查(NBS)是一个非常成功的公共卫生项目,其设计伴随着假阳性结果。在这里,我们训练了一个随机森林机器学习分类器来筛选数据,以提高对真假阳性的预测。数据包括串联质谱仪检测到的39种代谢分析物和临床变量,如孕周和出生体重。对加州国家统计局项目报告的2777例筛查阳性的队列进行了分析性能评估,其中包括235例确诊病例和2542例假阳性,涉及四种疾病之一:戊二酸血症1型(GA-1)、甲基丙二酸血症(MMA)、鸟氨酸转氨基甲酸酶缺乏症(OTCD)和超长链酰辅酶A脱氢酶缺乏症(VLCADD)。在不改变筛查中检测这些疾病的灵敏度的情况下,基于随机森林的所有代谢物分析将GA-1的假阳性数量减少了89%,MMA的假阳性减少了45%,OTCD的假阳性减少了98%,VLCADD的假阳性减少了2%。所有的主要疾病标记物和之前报道的分析物,如MMA和OTCD的蛋氨酸,都是排名最高的分析物。随机森林公司对GA-1假阳性进行分类的能力与使用临床实验室综合报告(CLIR)获得的结果相似。我们开发了一个在线随机森林工具,用于对新生儿筛查中日益复杂的数据进行解释性分析。
Newborn screening (NBS) for inborn metabolic disorders is a highly successful public health program that by design is accompanied by false-positive results. Here we trained a Random Forest machine learning classifier on screening data to improve prediction of true and false positives. Data included 39 metabolic analytes detected by tandem mass spectrometry and clinical variables such as gestational age and birth weight. Analytical performance was evaluated for a cohort of 2777 screen positives reported by the California NBS program, which consisted of 235 confirmed cases and 2542 false positives for one of four disorders: glutaric acidemia type 1 (GA-1), methylmalonic acidemia (MMA), ornithine transcarbamylase deficiency (OTCD), and very long-chain acyl-CoA dehydrogenase deficiency (VLCADD). Without changing the sensitivity to detect these disorders in screening, Random Forest-based analysis of all metabolites reduced the number of false positives for GA-1 by 89%, for MMA by 45%, for OTCD by 98%, and for VLCADD by 2%. All primary disease markers and previously reported analytes such as methionine for MMA and OTCD were among the top-ranked analytes. Random Forest's ability to classify GA-1 false positives was found similar to results obtained using Clinical Laboratory Integrated Reports (CLIR). We developed an online Random Forest tool for interpretive analysis of increasingly complex data from newborn screening.