Next Generation Phenotyping Using the Unified Medical Language System

Next Generation Phenotyping Using the Unified Medical Language System
复制标题

DOI:
10.2196/medinform.3172
复制
发表时间:
2014-01-01
影响因子:
3.2
通讯作者:
Shimoyama, Mary
Shimoyama, Mary
中科院分区:
医学3区
文献类型:
--
作者:
Adamusiak, Tomasz;Shimoyama, Naoki;Shimoyama, Mary

文献摘要

被引文献

相似文献

背景:患者病历中的结构化信息是一个很大程度上尚未开发的研究数据宝库。在美国,尽管存在隐私问题,但由于有意义的使用立法推动了越来越多的电子健康记录(EHR)和医疗保健数据标准的采用,这一点最近变得更容易获得。另一方面,现在越来越难驾驭大量不同的临床术语标准,这些标准往往跨越数百万个概念。目的:我们研究的目标是开发一种方法,用于集成大量结构化的临床信息,既不了解术语,又能够捕获不同的临床表型,包括问题、程序、药物和临床结果(如实验室测试和临床观察)。在此背景下,我们将表型定义为提取包含在EHR中的所有临床相关特征。方法:项目范围由公共有意义使用(MU)数据集术语标准、医学临床术语系统化术语(SNOMED CT)、RxNorm、逻辑观察标识符名称和代码(LOINC)、当前程序术语(CPT)、医疗保健通用程序编码系统(HCPCS)、国际疾病分类第九版临床修改(ICD-9-CM)和国际疾病分类第十版临床修改(ICD-10-CM)确定。使用统一医学语言系统(UMLS)作为MU本体之间的映射层。提取、加载和转换方法将EHR中的原始注释从映射过程中分离出来,并允许随着术语的更新而持续更新。此外,我们将所有术语集成到一个单一的UMLS派生本体中,并对其进行进一步优化,使相对较大的概念图更易于管理。结果:初步评估使用来自临床化身项目的模拟数据,使用了100,000名虚拟患者,他们接受了90天的、基因指导的华法林剂量方案。该数据集使用标准的MU术语进行注释,并使用UMLS进行加载和转换。我们使用在Froedtert医院接受治疗的7931名患者(1200万次临床观察)的结构化电子病历数据,在我们的内部分析平台中部署了此方法以进行规模扩展。使用凭据用户“jmirdemo”和密码“jmirdemo”,可以在互联网上获得仅限于临床化身数据的演示。结论:尽管UMLS固有的复杂性,但它可以作为当前医疗保健领域使用的许多临床数据标准的有效接口术语。
Background: Structured information within patient medical records represents a largely untapped treasure trove of research data. In the United States, privacy issues notwithstanding, this has recently become more accessible thanks to the increasing adoption of electronic health records (EHR) and health care data standards fueled by the Meaningful Use legislation. The other side of the coin is that it is now becoming increasingly more difficult to navigate the profusion of many disparate clinical terminology standards, which often span millions of concepts.Objective: The objective of our study was to develop a methodology for integrating large amounts of structured clinical information that is both terminology agnostic and able to capture heterogeneous clinical phenotypes including problems, procedures, medications, and clinical results (such as laboratory tests and clinical observations). In this context, we define phenotyping as the extraction of all clinically relevant features contained in the EHR.Methods: The scope of the project was framed by the Common Meaningful Use (MU) Dataset terminology standards; the Systematized Nomenclature of Medicine Clinical Terms (SNOMED CT), RxNorm, the Logical Observation Identifiers Names and Codes (LOINC), the Current Procedural Terminology (CPT), the Health care Common Procedure Coding System (HCPCS), the International Classification of Diseases Ninth Revision Clinical Modification (ICD-9-CM), and the International Classification of Diseases Tenth Revision Clinical Modification (ICD-10-CM). The Unified Medical Language System (UMLS) was used as a mapping layer among the MU ontologies. An extract, load, and transform approach separated original annotations in the EHR from the mapping process and allowed for continuous updates as the terminologies were updated. Additionally, we integrated all terminologies into a single UMLS derived ontology and further optimized it to make the relatively large concept graph manageable.Results: The initial evaluation was performed with simulated data from the Clinical Avatars project using 100,000 virtual patients undergoing a 90 day, genotype guided, warfarin dosing protocol. This dataset was annotated with standard MU terminologies, loaded, and transformed using the UMLS. We have deployed this methodology to scale in our in-house analytics platform using structured EHR data for 7931 patients (12 million clinical observations) treated at the Froedtert Hospital. A demonstration limited to Clinical Avatars data is available on the Internet using the credentials user "jmirdemo" and password "jmirdemo".Conclusions: Despite its inherent complexity, the UMLS can serve as an effective interface terminology for many of the clinical data standards currently used in the health care domain.