Machine Learning for the Biochemical Genetics Laboratory.

Machine Learning for the Biochemical Genetics Laboratory.
复制标题

DOI:
10.1093/clinchem/hvaa168
复制
发表时间:
2020-09-01
期刊:
影响因子:
9.3
通讯作者:
Master SR
Master SR
中科院分区:
医学1区
文献类型:
--
作者:
Ganetzky RD;Master SR

文献摘要

参考文献

被引文献

相似文献

机器学习已经成为现代数据分析不可或缺的一部分,特别是对于基于复杂变量集的样本分类。在临床诊断领域,它已成功地用于从识别扫描载玻片图像上的组织异常区域到基于实验室结果预测疾病进展的应用。简而言之,“机器学习”(人工智能的一个子集)从具有已知标签的样本数据集开始(在这种情况下,疾病诊断)。然后使用各种算法来“训练”分类器(基于该原始“训练集”),该分类器将预测先前未分析的新的未知样本中的适当诊断。通过在分析复杂的实验室数据集时提供计算“专业知识”,经验证的机器学习分类器可以作为临床实验室专业人员向一线临床医生提供解释或诊断指导的有价值的帮助。血浆氨基酸谱分析是诊断先天性代谢缺陷的最重要的检测方法之一。对氨基酸谱的解释可以很简单,例如从苯丙氨酸增加诊断苯丙酮尿症;然而,在其他情况下,细微的差别可能是唯一的诊断线索。氨基酸浓度的相互关联性难以量化用于诊断目的。虽然已经进行了研究来定义两种相关氨基酸之间的诊断比率(1),但这仅在最常见的氨基酸代谢疾病中进行了探索,并且大多数这些比率的参考范围不存在。根据氨基酸浓度进行准确和快速的诊断对于提供挽救生命的治疗至关重要,例如在尿素循环缺陷中。目前,从氨基酸建立诊断需要临床化学家或临床生化遗传学家的人工审查和临床解释。考虑到所涉及的细微差别,这种解释可能完全依赖于通过考虑氨基酸浓度的完整模式而发展出来的经验完形。提供这些口译的训练有素的工作人员有限。即使在熟练的手中,确定轻微的变化是否代表先天性代谢缺陷的亚型形式,或由于饮食摄入、药物使用引起的良性变化,或由于肝脏或肾脏功能障碍引起的继发性变化也是一个挑战。因此,这个面板是成熟的机器学习approaches.Previous研究计算支持,以提高临床灵敏度和特异性的氨基酸分析的应用程序已经在很大程度上是在新生儿筛查的背景下。有限的氨基酸定量是新生儿筛查的一个组成部分。然而,新生儿筛选氨基酸的挑战要有限得多。在新生儿筛查中仅测量氨基酸的一个子集,并且诊断集被限制为作为新生儿筛查程序的一部分的少数疾病。在这个空间内,用于确定分析物之间关系的计算分析后工具,在已知样品上训练,已经使用了几年(2)。最近,一种基于随机森林的机器学习方法已成功应用于新生儿筛查,包括氨基酸紊乱鸟氨酸转氨甲酰酶缺乏症(3)。在本期临床化学中,Wilkes等人提出了一种应用机器学习自动解释完整血浆氨基酸谱的方法(4)。与新生儿筛查期间分析的氨基酸相反,完整的氨基酸谱测量至少22种氨基酸,可用于分析新生儿的氨基酸。
Machine learning has emerged as an indispensable part of modern data analysis, particularly for classifying samples based on complex sets of variables. In the realm of clinical diagnostics, it has been successfully utilized for applications ranging from identifying abnormal areas of tissue on a scanned slide image, through predicting disease progression based on laboratory results. Briefly,“machine learning”(which is a subset of artificial intelligence) begins with a data set of samples with known labels (in this case, disease diagnoses). A variety of algorithms are then used to “train” a classifier (based on this original “training set”) that will predict the appropriate diagnosis in a new, unknown sample that has not been previously analyzed. By providing computational “expertise” in analyzing a complex laboratory data set, a validated machine learning classifier can serve as a valuable aid for the clinical laboratory professional in providing interpretive or diagnostic guidance to a front-line clinician. Plasma amino acid profiling is one of the most important tests for diagnosing inborn errors of metabolism. Interpretation of an amino acid profile can be straightforward, such as in diagnosing phenylketonuria from increased phenylalanine; however, in other cases, subtle nuances may be the only diagnostic clue. The interconnectedness of amino acid concentrations is difficult to quantify for diagnostic purposes. Although studies have been done to define diagnostic ratios between two related amino acids (1), this has only been explored in the most common disorders of amino acid metabolism, and reference ranges for most of these ratios do not exist. Accurate and rapid diagnosis from amino acid concentrations is essential for providing life-saving therapies, for example, in urea cycle defects. Currently, establishing a diagnosis from amino acids requires manual review and clinical interpretation by a clinical chemist or clinical biochemical geneticist. Given the nuanced variations involved, this interpretation may rely entirely on an experienced gestalt developed by considering the full pattern of amino acid concentrations. There is a limited trained workforce to provide these interpretations. Even in skilled hands, determining whether a slight variation represents a hypomorphic form of an inborn error of metabolism or benign variation due to dietary intake, drug use, or secondary variation due to liver or renal dysfunction is a challenge. Therefore, this panel is ripe for application of a machine learning approach.Previous research on computational support to improve the clinical sensitivity and specificity of amino acid analysis has largely been within the context of newborn screening. Limited amino acid quantification is an integral part of the newborn screen. However, the challenge of newborn screening amino acids is much more circumscribed. Only a subset of amino acids is measured in newborn screening, and the diagnostic set is constrained to a small number of disorders that are part of the newborn screening program. Within this space, computational postanalytical tools to determine relationships among analytes, trained on known samples, have been utilized for several years (2). More recently, a random forest-based machine learning approach has been successfully applied to newborn screening, including the amino acid disorder ornithine transcarbamylase deficiency (3). In this issue of Clinical Chemistry, Wilkes et al. present a method to apply machine learning to automate the interpretation of a full plasma amino acid profile (4). In contrast to amino acids analyzed during newborn screening, a full amino acid profile measures at least 22 amino acids and can be used to …
DOI: 10.3390/ijns6010016
发表时间: 2020-03-01
影响因子: 3.5
作者:
Peng, Gang;Tang, Yishuo;Scharfe, Curt
通讯作者: Scharfe, Curt
分析后工具改善了串联质谱法对新生儿筛查的性能。
DOI: 10.1038/gim.2014.62
发表时间: 2014-12
期刊: Genetics in medicine : official journal of the American College of Medical Genetics
影响因子: --
作者:
通讯作者: --
DOI: 10.1093/clinchem/hvaa134
发表时间: 2020-09-01
期刊: CLINICAL CHEMISTRY
影响因子: 9.3
作者:
Wilkes, Edmund H.;Emmett, Erin;Carling, Rachel S.
通讯作者: Carling, Rachel S.
DOI: 10.1007/8904_2012_186
发表时间: 2013-01-01
期刊: JIMD REPORTS - CASE AND RESEARCH REPORTS, 2012/6
影响因子: --
作者:
Rodney, S.;Boneh, A.
通讯作者: Boneh, A.