Hospital-wide Natural Language Processing summarising the health data of 1 million patients

Hospital-wide Natural Language Processing summarising the health data of 1 million patients
复制标题

全院自然语言处理汇总 100 万患者健康数据

DOI:
10.1101/2022.09.15.22279981
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Bean D
Bean D
中科院分区:
--
文献类型:
--
作者:
Bean D

文献摘要

相似文献

电子健康记录(EHR)代表真实的世界临床轨迹、干预和结果的主要存储库。虽然现代企业EHR试图以结构化标准化格式捕获数据,但EHR中捕获的大量可用信息仍然仅以非结构化文本格式记录,并且只能通过手动过程转换为结构化代码。最近,自然语言处理(NLP)算法已经达到了适合从临床文本中提取大规模和准确信息的性能水平。在这里,我们描述了应用程序的开源命名实体识别和链接(NER+L)的方法(CogStack,MedCAT)的整个文本内容的一个大型英国医院的信任(国王学院医院,伦敦)。由此产生的数据集包含1.57亿SNOMED概念,这些概念是从9年内107万患者的950万文档中生成的。我们总结了患病率和疾病发作,以及患者嵌入,捕捉大规模的主要合并症模式。NLP有潜力通过传统手动任务的大规模自动化来改变健康数据的生命周期。
Electronic health records (EHRs) represent a major repository of real world clinical trajectories, interventions and outcomes. While modern enterprise EHR’s try to capture data in structured standardised formats, a significant bulk of the available information captured in the EHR is still recorded only in unstructured text format and can only be transformed into structured codes by manual processes. Recently, Natural Language Processing (NLP) algorithms have reached a level of performance suitable for large scale and accurate information extraction from clinical text. Here we describe the application of open-source named-entity-recognition and linkage (NER+L) methods (CogStack, MedCAT) to the entire text content of a large UK hospital trust (King’s College Hospital, London). The resulting dataset contains 157M SNOMED concepts generated from 9.5M documents for 1.07M patients over a period of 9 years. We present a summary of prevalence and disease onset as well as a patient embedding that captures major comorbidity patterns at scale. NLP has the potential to transform the health data lifecycle, through large-scale automation of a traditionally manual task.