A large language model for electronic health records.

A large language model for electronic health records.
复制标题

电子健康记录的大型语言模型。

DOI:
10.1038/s41746-022-00742-2
复制
发表时间:
2022-12-26
影响因子:
15.2
通讯作者:
--
中科院分区:
医学1区
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

人们对开发人工智能(AI)系统来处理和解释电子健康记录(EHRs)越来越感兴趣。基于预训练语言模型的自然语言处理(NLP)是利用临床叙述的医疗人工智能系统的关键技术。然而,临床语言模型很少,其中在临床领域训练的最大的模型相对较小,只有1.1亿个参数(与一般领域的数十亿个参数相比)。目前尚不清楚拥有数十亿参数的大型临床语言模型如何帮助医疗人工智能系统利用非结构化电子病历。在这项研究中,我们从无开始开发了一个大型临床语言模型gatortron,该模型使用了900亿字的文本(包括820亿字的去识别临床文本),并在临床概念提取、医学关系提取、语义文本相似度、自然语言推理(NLI)和医学问答(MQA)五个临床NLP任务上对其进行了系统的评估。我们研究了(1)扩大参数数量和(2)扩大训练数据的大小如何有利于这些NLP任务。GatorTron模型将临床语言模型从1.1亿个参数扩展到89亿个参数,并改进了五个临床NLP任务(例如,NLI和MQA的准确性分别提高了9.6%和9.5%),这可以应用于医疗人工智能系统,以改善医疗服务。GatorTron模型可以在https://catalog.ngc.nvidia.com/orgs/nvidia/teams/clara/models/gatortron_og上公开获取。
There is an increasing interest in developing artificial intelligence (AI) systems to process and interpret electronic health records (EHRs). Natural language processing (NLP) powered by pretrained language models is the key technology for medical AI systems utilizing clinical narratives. However, there are few clinical language models, the largest of which trained in the clinical domain is comparatively small at 110 million parameters (compared with billions of parameters in the general domain). It is not clear how large clinical language models with billions of parameters can help medical AI systems utilize unstructured EHRs. In this study, we develop from scratch a large clinical language model—GatorTron—using >90 billion words of text (including >82 billion words of de-identified clinical text) and systematically evaluate it on five clinical NLP tasks including clinical concept extraction, medical relation extraction, semantic textual similarity, natural language inference (NLI), and medical question answering (MQA). We examine how (1) scaling up the number of parameters and (2) scaling up the size of the training data could benefit these NLP tasks. GatorTron models scale up the clinical language model from 110 million to 8.9 billion parameters and improve five clinical NLP tasks (e.g., 9.6% and 9.5% improvement in accuracy for NLI and MQA), which can be applied to medical AI systems to improve healthcare delivery. The GatorTron models are publicly available at: https://catalog.ngc.nvidia.com/orgs/nvidia/teams/clara/models/gatortron_og.
DOI: 10.1007/s10916-017-0716-5
发表时间: 2017-05-01
影响因子: 5.3
作者:
Bush, Ruth A.;Kuelbs, Cynthia;Chiang, George
通讯作者: Chiang, George
DOI: 10.1038/sdata.2016.35
发表时间: 2016-05-24
期刊: Scientific data
影响因子: 9.8
作者:
Johnson AE;Pollard TJ;Shen L;Lehman LW;Feng M;Ghassemi M;Moody B;Szolovits P;Celi LA;Mark RG
通讯作者: Mark RG
DOI: 10.1093/bioinformatics/btz682
发表时间: 2020-02-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Lee J;Yoon W;Kim S;Kim D;Kim S;So CH;Kang J
通讯作者: Kang J
DOI: 10.1093/jamia/ocx080
发表时间: 2017-11-01
影响因子: 6.4
作者:
Adler-Milstein, Julia;Holmgren, A. Jay;Patel, Vaishali
通讯作者: Patel, Vaishali
利用人工智能对儿科疾病进行评估和准确诊断
DOI: 10.1038/s41591-018-0335-9
发表时间: 2019-03-01
期刊: NATURE MEDICINE
影响因子: 82.9
作者:
Liang, Huiying;Tsui, Brian Y.;Xia, Huimin
通讯作者: Xia, Huimin