sureLDA: A multidisease automated phenotyping method for the electronic health record

sureLDA: A multidisease automated phenotyping method for the electronic health record
复制标题

DOI:
10.1093/jamia/ocaa079
复制
发表时间:
2020-08-01
影响因子:
6.4
通讯作者:
Cai, Tianxi
Cai, Tianxi
中科院分区:
管理学2区
文献类型:
--
作者:
Ahuja, Yuri;Zhou, Doudou;Cai, Tianxi

文献摘要

被引文献

相似文献

目的:阻碍电子健康记录数据用于翻译研究的一个主要瓶颈是缺乏准确的表型标签。图表审查以及基于规则和监督的表型方法需要费力的专家输入,阻碍了对需要定义和标记从头开始的许多表型的研究的适用性。虽然在这种情况下,国际疾病分类代码经常被用作真实标签的替代品,但这些代码有时缺乏特异性。我们提出了一种全自动的主题建模算法来同时标注多个表型。材料和方法:代理引导的集成潜在狄利克雷分配(SureLDA)是一种无标签的多维表型识别方法。它首先使用PheNorm算法基于每个目标表型的2个代理特征初始化概率,然后利用这些概率来约束LDA主题模型以生成特定于表型的主题。最后,它通过聚类集成将表型特征计数与替代项相结合,以产生最终的表型概率。结果:确保LDA在一系列模拟和真实的表型上实现了可靠的高准确度和精确度。它的性能对替代和非替代特征的表型流行率和相对信息量是稳健的。讨论:当然,LDA结合了PheNorm和LDA的吸引人的特性,实现了对不同表型特征的高精度和精确度。它为一些代理功能不能充分捕获的表型提供了特别的改进。此外,Sure LDA的特征选择能力使其能够处理高特征维度并产生可解释的计算表型。结论:Sure LDA非常适合于大规模电子健康记录表型分析,用于高度多表型应用,如表型全组研究。
Objective: A major bottleneck hindering utilization of electronic health record data for translational research is the lack of precise phenotype labels. Chart review as well as rule-based and supervised phenotyping approaches require laborious expert input, hampering applicability to studies that require many phenotypes to be defined and labeled de novo. Though International Classification of Diseases codes are often used as surrogates for true labels in this setting, these sometimes suffer from poor specificity. We propose a fully automated topic modeling algorithm to simultaneously annotate multiple phenotypes.Materials and Methods: Surrogate-guided ensemble latent Dirichlet allocation (sureLDA) is a label-free multidimensional phenotyping method. It first uses the PheNorm algorithm to initialize probabilities based on 2 surrogate features for each target phenotype, and then leverages these probabilities to constrain the LDA topic model to generate phenotype-specific topics. Finally, it combines phenotype-feature counts with surrogates via clustering ensemble to yield final phenotype probabilities.Results: sureLDA achieves reliably high accuracy and precision across a range of simulated and real-world phenotypes. Its performance is robust to phenotype prevalence and relative informativeness of surogate vs nonsurrogate features. It also exhibits powerful feature selection properties.Discussion: sureLDA combines attractive properties of PheNorm and LDA to achieve high accuracy and precision robust to diverse phenotype characteristics. It offers particular improvement for phenotypes insufficiently captured by a few surrogate features. Moreover, sureLDA's feature selection ability enables it to handle high feature dimensions and produce interpretable computational phenotypes.Conclusions: sureLDA is well suited toward large-scale electronic health record phenotyping for highly multiphenotype applications such as phenome-wide association studies.