Hospital-wide natural language processing summarising the health data of 1 million patients.

Hospital-wide natural language processing summarising the health data of 1 million patients.
复制标题

DOI:
10.1371/journal.pdig.0000218
复制
发表时间:
2023-05
期刊:
PLOS digital health
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

相似文献

电子健康记录(EHRs)是真实世界临床轨迹、干预措施和结果的主要存储库。虽然现代企业EHR试图以结构化的标准化格式捕获数据,但在EHR中捕获的大量可用信息仍然仅以非结构化文本格式记录,并且只能通过手动过程转换为结构化代码。近年来,自然语言处理(NLP)算法已经达到了适合大规模、准确提取临床文本信息的性能水平。在这里,我们描述了开源命名实体识别和链接(NER+L)方法(CogStack, MedCAT)在英国一家大型医院信托基金(伦敦国王学院医院)的整个文本内容中的应用。由此产生的数据集包含157M个SNOMED概念,这些概念是在9年的时间里从950万份文档中为107万名患者生成的。我们提出了患病率和疾病发病的总结,以及一个病人嵌入捕获主要合并症模式的规模。NLP通过传统手工任务的大规模自动化,有可能改变健康数据的生命周期。临床笔记和信函仍然是临床人员记录和共享医疗信息的主要方式。这意味着对于研究,我们需要能够处理文本数据的方法,这通常比诊断代码或测试结果等“结构化”数据更具挑战性。在这项研究中,我们应用最先进的临床文本处理模型来分析伦敦一家大型医院近10年的文本数据,涵盖超过100万名患者。我们能够在纯文本数据中找到疾病负担、发病和共发的模式。这一结果有力地支持了临床文献数据在研究中的使用,并为其他研究人员提供了临床文献规模和性质的总结。
Electronic health records (EHRs) represent a major repository of real world clinical trajectories, interventions and outcomes. While modern enterprise EHR’s try to capture data in structured standardised formats, a significant bulk of the available information captured in the EHR is still recorded only in unstructured text format and can only be transformed into structured codes by manual processes. Recently, Natural Language Processing (NLP) algorithms have reached a level of performance suitable for large scale and accurate information extraction from clinical text. Here we describe the application of open-source named-entity-recognition and linkage (NER+L) methods (CogStack, MedCAT) to the entire text content of a large UK hospital trust (King’s College Hospital, London). The resulting dataset contains 157M SNOMED concepts generated from 9.5M documents for 1.07M patients over a period of 9 years. We present a summary of prevalence and disease onset as well as a patient embedding that captures major comorbidity patterns at scale. NLP has the potential to transform the health data lifecycle, through large-scale automation of a traditionally manual task. Clinical notes and letters are still the main way that medical information is recorded and shared between clinical staff. This means that for research we need methods that can cope with text data, which is typically far more challenging than “structured” data like diagnosis codes or test results. In this study we apply a state of the art clinical text processing model to analyse almost 10 years worth of text data from a large London hospital, covering over 1 million patients. We are able to find patterns of disease burden, onset and co-occurrence purely in text data. This result strongly supports the use of clinical text data in research and provides a summary of the scale and nature of clinical text to other researchers.