Med-BERT: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction.
Med-BERT: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction.
复制标题
Med-BERT:在大规模结构化电子健康记录上进行预训练的情境化嵌入,用于疾病预测。
DOI:
10.1038/s41746-021-00455-y
复制
发表时间:
2021-05-20
影响因子:
15.2
通讯作者:
Zhi D
中科院分区:
文献类型:
--
作者:
Rasmy L;Xiang Y;Xie Z;Tao C;Zhi D
Deep learning (DL)-based predictive models from electronic health records (EHRs) deliver impressive performance in many clinical tasks. Large training cohorts, however, are often required by these models to achieve high accuracy, hindering the adoption of DL-based models in scenarios with limited training data. Recently, bidirectional encoder representations from transformers (BERT) and related models have achieved tremendous successes in the natural language processing domain. The pretraining of BERT on a very large training corpus generates contextualized embeddings that can boost the performance of models trained on smaller datasets. Inspired by BERT, we propose Med-BERT, which adapts the BERT framework originally developed for the text domain to the structured EHR domain. Med-BERT is a contextualized embedding model pretrained on a structured EHR dataset of 28,490,650 patients. Fine-tuning experiments showed that Med-BERT substantially improves the prediction accuracy, boosting the area under the receiver operating characteristics curve (AUC) by 1.21–6.14% in two disease prediction tasks from two clinical databases. In particular, pretrained Med-BERT obtains promising performances on tasks with small fine-tuning training sets and can boost the AUC by more than 20% or obtain an AUC as high as a model trained on a training set ten times larger, compared with deep learning models without Med-BERT. We believe that Med-BERT will benefit disease prediction studies with small local training datasets, reduce data collection expenses, and accelerate the pace of artificial intelligence aided healthcare.
登录
查看更多内容
影响因子:
64.8
作者:
Esteva A;Kuprel B;Novoa RA;Ko J;Swetter SM;Blau HM;Thrun S
通讯作者:
Thrun S
影响因子:
9.8
作者:
Johnson AE;Pollard TJ;Shen L;Lehman LW;Feng M;Ghassemi M;Moody B;Szolovits P;Celi LA;Mark RG
通讯作者:
Mark RG
影响因子:
5.9
作者:
Jiang F;Jiang Y;Zhi H;Dong Y;Li H;Ma S;Wang Y;Dong Q;Shen H;Wang Y
通讯作者:
Wang Y
影响因子:
5.9
作者:
Gupta, Priyanka;Malhotra, Pankaj;Shroff, Gautam
通讯作者:
Shroff, Gautam
影响因子:
3.5
作者:
Gentil ML;Cuggia M;Fiquet L;Hagenbourger C;Le Berre T;Banâtre A;Renault E;Bouzille G;Chapron A
通讯作者:
Chapron A