Development and application of a high throughput natural language processing architecture to convert all clinical documents in a clinical data warehouse into standardized medical vocabularies

Development and application of a high throughput natural language processing architecture to convert all clinical documents in a clinical data warehouse into standardized medical vocabularies
复制标题

DOI:
10.1093/jamia/ocz068
复制
发表时间:
2019-11-01
影响因子:
6.4
通讯作者:
Price, Ron
Price, Ron
中科院分区:
管理学2区
文献类型:
--
作者:
Afshar, Majid;Dligach, Dmitriy;Price, Ron

文献摘要

被引文献

相似文献

目的:自然语言处理(NLP)引擎,如临床文本分析和知识提取系统,是处理研究笔记的解决方案,但优化其临床数据仓库的性能仍然是一个挑战。我们的目标是开发一个高通量的NLP架构,使用临床文本分析和知识提取系统,并提出了一个预测模型usecase.Materials和方法:CDW是由1 103 038名患者在10年。该架构使用Hadoop数据存储库构建源数据,并使用3个大规模对称处理服务器构建NLP。临床文档中提到的每个命名实体都映射到统一医学语言系统概念唯一标识符(CUI)。结果:NLP架构在13.33天内处理了83 867 802份临床文档,并在8个标准化医学词汇表中产生了37 721 886 606个CUI。该架构的性能超过500 000文档每小时在30个并行实例的临床文本分析和知识提取系统,包括10个实例专用于文件大于20 000字节。在预测30天再入院的用例示例中,基于CUI的模型与n-gram具有相似的区分度,曲线下面积接收器操作特征为0.75(95%CI,0.74-0.76)。讨论和结论:我们的卫生系统的高通量NLP架构可以作为使用基于CUI的方法进行大规模临床研究的基准。
Objective: Natural language processing (NLP) engines such as the clinical Text Analysis and Knowledge Extraction System are a solution for processing notes for research, but optimizing their performance for a clinical data warehouse remains a challenge. We aim to develop a high throughput NLP architecture using the clinical Text Analysis and Knowledge Extraction System and present a predictive model use case.Materials and Methods: The CDW was comprised of 1 103 038 patients across 10 years. The architecture was constructed using the Hadoop data repository for source data and 3 large-scale symmetric processing servers for NLP. Each named entity mention in a clinical document was mapped to the Unified Medical Language System concept unique identifier (CUI).Results: The NLP architecture processed 83 867 802 clinical documents in 13.33 days and produced 37 721 886 606 CUIs across 8 standardized medical vocabularies. Performance of the architecture exceeded 500 000 documents per hour across 30 parallel instances of the clinical Text Analysis and Knowledge Extraction System including 10 instances dedicated to documents greater than 20 000 bytes. In a use-case example for predicting 30-day hospital readmission, a CUI-based model had similar discrimination to n-grams with an area under the curve receiver operating characteristic of 0.75 (95% CI, 0.74-0.76).Discussion and Conclusion: Our health system's high throughput NLP architecture may serve as a benchmark for large-scale clinical research using a CUI-based approach.