Discovering Patient Phenotypes Using Generalized Low Rank Models

Discovering Patient Phenotypes Using Generalized Low Rank Models
复制标题

DOI:
10.1142/9789814749411_0014
复制
发表时间:
2016
影响因子:
--
通讯作者:
Alejandro Schuler;V. Liu;J. Wan;A. Callahan;Madeleine Udell;David E. Stark;N. Shah
Alejandro Schuler;V. Liu;J. Wan;A. Callahan;Madeleine Udell;David E. Stark;N. Shah
中科院分区:
--
文献类型:
--
作者:
Alejandro Schuler;V. Liu;J. Wan;A. Callahan;Madeleine Udell;David E. Stark;N. Shah

文献摘要

相似文献

医学实践的前提是发现患者之间的共性或区别特征,以告知相应的治疗。给定患者分组(下文称为表型),临床医生可以实施说明该表型中疾病的根本原因的治疗途径。传统上,表型是通过直觉、实践经验和基础科学的进步发现的,但这些方法通常是启发式的、劳动密集型的,可能需要几十年才能产生可操作的知识。尽管在过去的世纪里,我们对疾病的理解有了很大的进步,但仍有一些重要的领域我们的表型是模糊的,比如在行为健康或医院环境中。为了加速表型发现,研究人员使用机器学习来寻找电子健康记录中的模式,但经常受到数据缺失、稀疏性和数据异质性的阻碍。在这项研究中,我们使用了一个灵活的框架,称为广义低秩模型(GLRM),以克服这些障碍,并发现表型在两个来源的患者数据。首先,我们分析了来自2010年医疗保健成本和利用项目国家住院患者样本(NIS)的数据,其中包含超过800万条住院记录,包括行政代码和人口统计信息。其次,我们分析了一个小的(N=1746),本地数据集记录自闭症谱系障碍患者的临床进展使用颗粒特征的电子健康记录,包括从医生笔记的文本。我们证明,低秩建模成功地捕获了这些截然不同的数据集中已知和推定的表型。
The practice of medicine is predicated on discovering commonalities or distinguishing characteristics among patients to inform corresponding treatment. Given a patient grouping (hereafter referred to as a phenotype), clinicians can implement a treatment pathway accounting for the underlying cause of disease in that phenotype. Traditionally, phenotypes have been discovered by intuition, experience in practice, and advancements in basic science, but these approaches are often heuristic, labor intensive, and can take decades to produce actionable knowledge. Although our understanding of disease has progressed substantially in the past century, there are still important domains in which our phenotypes are murky, such as in behavioral health or in hospital settings. To accelerate phenotype discovery, researchers have used machine learning to find patterns in electronic health records, but have often been thwarted by missing data, sparsity, and data heterogeneity. In this study, we use a flexible framework called Generalized Low Rank Modeling (GLRM) to overcome these barriers and discover phenotypes in two sources of patient data. First, we analyze data from the 2010 Healthcare Cost and Utilization Project National Inpatient Sample (NIS), which contains upwards of 8 million hospitalization records consisting of administrative codes and demographic information. Second, we analyze a small (N=1746), local dataset documenting the clinical progression of autism spectrum disorder patients using granular features from the electronic health record, including text from physician notes. We demonstrate that low rank modeling successfully captures known and putative phenotypes in these vastly different datasets.