Feature extraction for phenotyping from semantic and knowledge resources

Feature extraction for phenotyping from semantic and knowledge resources
复制标题

从语义和知识资源中提取表型特征

DOI:
10.1016/j.jbi.2019.103122
复制
发表时间:
2019-03-01
影响因子:
4.5
通讯作者:
Yu, Sheng
Yu, Sheng
中科院分区:
医学3区
文献类型:
--
作者:
Ning, Wenxin;Chan, Stephanie;Yu, Sheng

文献摘要

被引文献

相似文献

目的:表型分型算法可以有效准确地识别具有特定疾病表型的患者,并构建基于电子健康记录(EHR)的队列,用于后续的临床或基因组研究。以前的研究已经引入了无监督的基于EHR的特征选择方法,这些方法产生了高精度的算法。然而,这些选择方法仍然需要专家干预,以根据每个表型的EHR数据分布调整参数设置。为了进一步加速表型分析算法的发展,我们提出了一种完全自动化和鲁棒的无监督特征选择方法,该方法仅利用公开可用的医学知识来源,而不是EHR数据。SEmantics-Driven Feature Extraction(SEDFE)从在线知识源中收集医学概念作为候选特征,并将其向量化。形成由神经词嵌入和统一医学语言系统元词库导出的分布式语义表示。通过线性分解标准确定语义上最接近并且充分表征目标表型的多个特征,并选择这些特征用于最终的分类算法。将SEDFE与基于EHR的SAFE算法和领域专家在特征选择方面进行了比较,用于五种表型的分类,包括冠状动脉疾病,类风湿性关节炎,克罗恩病,溃疡性结肠炎,和小儿肺动脉高压的研究。SEDFE产生的算法实现了与SAFE和expertcurated功能产生的算法相当的准确性。结论:SEDFE在EHR表型的无监督特征选择中取得了令人满意的性能。无论是完全自动化和EHR的独立,这种方法的效率和准确性,在开发算法的高通量表型。
Objective: Phenotyping algorithms can efficiently and accurately identify patients with a specific disease phenotype and construct electronic health records (EHR)-based cohorts for subsequent clinical or genomic studies. Previous studies have introduced unsupervised EHR-based feature selection methods that yielded algorithms with high accuracy. However, those selection methods still require expert intervention to tweak the parameter settings according to the EHR data distribution for each phenotype. To further accelerate the development of phenotyping algorithms, we propose a fully automated and robust unsupervised feature selection method that leverages only publicly available medical knowledge sources, instead of EHR data.Methods: SEmantics-Driven Feature Extraction (SEDFE) collects medical concepts from online knowledge sources as candidate features and gives them vector-form distributional semantic representations derived with neural word embedding and the Unified Medical Language System Metathesaurus. A number of features that are semantically closest and that sufficiently characterize the target phenotype are determined by a linear decomposition criterion and are selected for the final classification algorithm.Results: SEDFE was compared with the EHR-based SAFE algorithm and domain experts on feature selection for the classification of five phenotypes including coronary artery disease, rheumatoid arthritis, Crohn's disease, ulcerative colitis, and pediatric pulmonary arterial hypertension using both supervised and unsupervised approaches. Algorithms yielded by SEDFE achieved comparable accuracy to those yielded by SAFE and expertcurated features. SEDFE is also robust to the input semantic vectors.Conclusion: SEDFE attains satisfying performance in unsupervised feature selection for EHR phenotyping. Both fully automated and EHR-independent, this method promises efficiency and accuracy in developing algorithms for high-throughput phenotyping.