Med-BERT: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction.

Med-BERT: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction.
复制标题

Med-BERT:在大规模结构化电子健康记录上进行预训练的情境化嵌入,用于疾病预测。

DOI:
10.1038/s41746-021-00455-y
复制
发表时间:
2021-05-20
影响因子:
15.2
通讯作者:
Zhi D
Zhi D
中科院分区:
医学1区
文献类型:
--
作者:
Rasmy L;Xiang Y;Xie Z;Tao C;Zhi D

文献摘要

参考文献

被引文献

相似文献

基于深度学习(DL)的电子健康记录(EHR)预测模型在许多临床任务中表现出色。然而,这些模型通常需要大量的训练队列来实现高精度,这阻碍了在训练数据有限的情况下采用基于DL的模型。最近,双向编码器表示从变压器(BERT)和相关的模型已经取得了巨大的成功,在自然语言处理领域。BERT在一个非常大的训练语料库上的预训练生成了上下文化的嵌入,可以提高在较小数据集上训练的模型的性能。受BERT的启发,我们提出了Med-BERT,它将最初为文本域开发的BERT框架适应于结构化EHR域。Med-BERT是一个在28,490,650名患者的结构化EHR数据集上预训练的上下文嵌入模型。微调实验表明,Med-BERT大大提高了预测精度,在两个临床数据库的两个疾病预测任务中,将受试者工作特征曲线下面积(AUC)提高了1.21-6.14%。特别是,与没有Med-BERT的深度学习模型相比,预训练的Med-BERT在使用小的微调训练集的任务上获得了有希望的性能,并且可以将AUC提高20%以上,或者获得与在训练集上训练的模型一样高的AUC。我们相信Med-BERT将有利于使用小型本地训练数据集进行疾病预测研究,降低数据收集费用,并加快人工智能辅助医疗保健的步伐。
Deep learning (DL)-based predictive models from electronic health records (EHRs) deliver impressive performance in many clinical tasks. Large training cohorts, however, are often required by these models to achieve high accuracy, hindering the adoption of DL-based models in scenarios with limited training data. Recently, bidirectional encoder representations from transformers (BERT) and related models have achieved tremendous successes in the natural language processing domain. The pretraining of BERT on a very large training corpus generates contextualized embeddings that can boost the performance of models trained on smaller datasets. Inspired by BERT, we propose Med-BERT, which adapts the BERT framework originally developed for the text domain to the structured EHR domain. Med-BERT is a contextualized embedding model pretrained on a structured EHR dataset of 28,490,650 patients. Fine-tuning experiments showed that Med-BERT substantially improves the prediction accuracy, boosting the area under the receiver operating characteristics curve (AUC) by 1.21–6.14% in two disease prediction tasks from two clinical databases. In particular, pretrained Med-BERT obtains promising performances on tasks with small fine-tuning training sets and can boost the AUC by more than 20% or obtain an AUC as high as a model trained on a training set ten times larger, compared with deep learning models without Med-BERT. We believe that Med-BERT will benefit disease prediction studies with small local training datasets, reduce data collection expenses, and accelerate the pace of artificial intelligence aided healthcare.
DOI: 10.1038/nature21056
发表时间: 2017-02-02
期刊: Nature
影响因子: 64.8
作者:
Esteva A;Kuprel B;Novoa RA;Ko J;Swetter SM;Blau HM;Thrun S
通讯作者: Thrun S
DOI: 10.1038/sdata.2016.35
发表时间: 2016-05-24
期刊: Scientific data
影响因子: 9.8
作者:
Johnson AE;Pollard TJ;Shen L;Lehman LW;Feng M;Ghassemi M;Moody B;Szolovits P;Celi LA;Mark RG
通讯作者: Mark RG
DOI: 10.1136/svn-2017-000101
发表时间: 2017-12
影响因子: 5.9
作者:
Jiang F;Jiang Y;Zhi H;Dong Y;Li H;Ma S;Wang Y;Dong Q;Shen H;Wang Y
通讯作者: Wang Y
DOI: 10.1007/s41666-019-00062-3
发表时间: 2020-06-01
影响因子: 5.9
作者:
Gupta, Priyanka;Malhotra, Pankaj;Shroff, Gautam
通讯作者: Shroff, Gautam
DOI: 10.1186/s12911-017-0538-x
发表时间: 2017-09-25
影响因子: 3.5
作者:
Gentil ML;Cuggia M;Fiquet L;Hagenbourger C;Le Berre T;Banâtre A;Renault E;Bouzille G;Chapron A
通讯作者: Chapron A