A clinical specific BERT developed using a huge Japanese clinical text corpus.
A clinical specific BERT developed using a huge Japanese clinical text corpus.
复制标题
DOI:
10.1371/journal.pone.0259763
复制
发表时间:
2021
期刊:
影响因子:
3.7
通讯作者:
Ohe K
中科院分区:
文献类型:
--
作者:
Kawazoe Y;Shibata D;Shinohara E;Aramaki E;Ohe K
Generalized language models that are pre-trained with a large corpus have achieved great performance on natural language tasks. While many pre-trained transformers for English are published, few models are available for Japanese text, especially in clinical medicine. In this work, we demonstrate the development of a clinical specific BERT model with a huge amount of Japanese clinical text and evaluate it on the NTCIR-13 MedWeb that has fake Twitter messages regarding medical concerns with eight labels. Approximately 120 million clinical texts stored at the University of Tokyo Hospital were used as our dataset. The BERT-base was pre-trained using the entire dataset and a vocabulary including 25,000 tokens. The pre-training was almost saturated at about 4 epochs, and the accuracies of Masked-LM and Next Sentence Prediction were 0.773 and 0.975, respectively. The developed BERT did not show significantly higher performance on the MedWeb task than the other BERT models that were pre-trained with Japanese Wikipedia text. The advantage of pre-training on clinical text may become apparent in more complex tasks on actual clinical text, and such an evaluation set needs to be developed.
影响因子:
4.8
作者:
Lampiris, Georgios;Karelakis, Christos;Loizou, Efstratios
通讯作者:
Loizou, Efstratios
DOI:
10.1044/2019_ajslp-cac48-18-0220
发表时间:
2020-02-01
影响因子:
2.6
作者:
Lee, Jaime B.;Azios, Jamie H.
通讯作者:
Azios, Jamie H.