Implementation of a Cohort Retrieval System for Clinical Data Repositories Using the Observational Medical Outcomes Partnership Common Data Model: Proof-of-Concept System Validation.

Implementation of a Cohort Retrieval System for Clinical Data Repositories Using the Observational Medical Outcomes Partnership Common Data Model: Proof-of-Concept System Validation.
复制标题

DOI:
10.2196/17376
复制
发表时间:
2020-10-06
影响因子:
3.2
通讯作者:
Liu H
Liu H
中科院分区:
医学3区
文献类型:
--
作者:
Liu S;Wang Y;Wen A;Wang L;Hong N;Shen F;Bedrick S;Hersh W;Liu H

文献摘要

参考文献

被引文献

相似文献

电子健康记录的广泛采用使得电子健康记录数据能够用于临床研究和医疗保健提供的二次使用。自然语言处理技术在其提取嵌入在非结构化临床数据中的信息的能力方面显示出了希望,并且信息检索技术提供了灵活且可扩展的解决方案,可以增强用于检索和排名相关记录的自然语言处理系统。在本文中,我们提出了一个队列检索系统,可以执行文本队列选择查询的结构化数据和非结构化的文本队列检索增强分析文本从电子健康记录(CREATE)的实现。CREATE是一个概念验证系统,它利用结构化查询和信息检索技术对自然语言处理结果的组合,使用观察性医学成果伙伴关系通用数据模型提高队列检索性能,以增强模型的可移植性。自然语言处理组件用于从文本查询中提取公共数据模型概念。我们设计了一个层次索引,以支持公共数据模型的概念搜索,利用信息检索技术和框架。我们对5个队列识别查询的案例研究,在患者级别和文档级别使用5个信息检索指标的精度进行评估,表明CREATE的平均精度为0.90,优于仅使用结构化数据或仅使用非结构化文本的系统,平均精度分别为0.54和0.74。马约诊所生物库数据的实施和评估表明,CREATE优于队列检索系统,在复杂的文本队列查询中只使用结构化数据或非结构化文本之一。
Widespread adoption of electronic health records has enabled the secondary use of electronic health record data for clinical research and health care delivery. Natural language processing techniques have shown promise in their capability to extract the information embedded in unstructured clinical data, and information retrieval techniques provide flexible and scalable solutions that can augment natural language processing systems for retrieving and ranking relevant records. In this paper, we present the implementation of a cohort retrieval system that can execute textual cohort selection queries on both structured data and unstructured text—Cohort Retrieval Enhanced by Analysis of Text from Electronic Health Records (CREATE). CREATE is a proof-of-concept system that leverages a combination of structured queries and information retrieval techniques on natural language processing results to improve cohort retrieval performance using the Observational Medical Outcomes Partnership Common Data Model to enhance model portability. The natural language processing component was used to extract common data model concepts from textual queries. We designed a hierarchical index to support the common data model concept search utilizing information retrieval techniques and frameworks. Our case study on 5 cohort identification queries, evaluated using the precision at 5 information retrieval metric at both the patient-level and document-level, demonstrates that CREATE achieves a mean precision at 5 of 0.90, which outperforms systems using only structured data or only unstructured text with mean precision at 5 values of 0.54 and 0.74, respectively. The implementation and evaluation of Mayo Clinic Biobank data demonstrated that CREATE outperforms cohort retrieval systems that only use one of either structured data or unstructured text in complex textual cohort queries.
DOI: 10.1038/s41586-018-0579-z
发表时间: 2018-10
期刊: Nature
影响因子: 64.8
作者:
Bycroft C;Freeman C;Petkova D;Band G;Elliott LT;Sharp K;Motyer A;Vukcevic D;Delaneau O;O'Connell J;Cortes A;Welsh S;Young A;Effingham M;McVean G;Leslie S;Allen N;Donnelly P;Marchini J
通讯作者: Marchini J
DOI: 10.1136/jamia.2009.001560
发表时间: 2010-09-01
影响因子: 6.4
作者:
Savova, Guergana K.;Masanz, James J.;Chute, Christopher G.
通讯作者: Chute, Christopher G.
DOI: 10.1016/j.jbi.2015.09.010
发表时间: 2015-12
影响因子: 4.5
作者:
Chen Y;Lasko TA;Mei Q;Denny JC;Xu H
通讯作者: Xu H
DOI: 10.1093/jamia/ocx019
发表时间: 2017-11-01
影响因子: 6.4
作者:
Kang, Tian;Zhang, Shaodian;Weng, Chunhua
通讯作者: Weng, Chunhua
DOI: 10.1001/jama.2011.1204
发表时间: 2011-08-24
影响因子: 120.7
作者:
Murff, Harvey J.;FitzHenry, Fern;Speroff, Theodore
通讯作者: Speroff, Theodore