A clinical specific BERT developed using a huge Japanese clinical text corpus.

A clinical specific BERT developed using a huge Japanese clinical text corpus.
复制标题

DOI:
10.1371/journal.pone.0259763
复制
发表时间:
2021
期刊:
影响因子:
3.7
通讯作者:
Ohe K
Ohe K
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Kawazoe Y;Shibata D;Shinohara E;Aramaki E;Ohe K

文献摘要

参考文献

被引文献

相似文献

使用大型语料库进行预训练的广义语言模型在自然语言任务中取得了很好的效果。虽然已经出版了许多针对英语的预训练转换器,但很少有模型可用于日语文本,特别是在临床医学中。在这项工作中,我们展示了一个具有大量日本临床文本的临床特异性BERT模型的开发,并在ntir -13 MedWeb上对其进行了评估,该MedWeb具有关于医疗问题的假Twitter消息,并带有八个标签。存储在东京大学医院的大约1.2亿临床文本被用作我们的数据集。bert库使用整个数据集和包含25,000个令牌的词汇表进行预训练。预训练在约4个epoch时基本饱和,mask - lm和Next Sentence Prediction的准确率分别为0.773和0.975。开发的BERT在MedWeb任务上的表现并不比其他用日语维基百科文本预训练的BERT模型高得多。临床文本预训练的优势可能会在更复杂的实际临床文本任务中显现出来,这样的评估集需要开发。
Generalized language models that are pre-trained with a large corpus have achieved great performance on natural language tasks. While many pre-trained transformers for English are published, few models are available for Japanese text, especially in clinical medicine. In this work, we demonstrate the development of a clinical specific BERT model with a huge amount of Japanese clinical text and evaluate it on the NTCIR-13 MedWeb that has fake Twitter messages regarding medical concerns with eight labels. Approximately 120 million clinical texts stored at the University of Tokyo Hospital were used as our dataset. The BERT-base was pre-trained using the entire dataset and a vocabulary including 25,000 tokens. The pre-training was almost saturated at about 4 epochs, and the accuracies of Masked-LM and Next Sentence Prediction were 0.773 and 0.975, respectively. The developed BERT did not show significantly higher performance on the MedWeb task than the other BERT models that were pre-trained with Japanese Wikipedia text. The advantage of pre-training on clinical text may become apparent in more complex tasks on actual clinical text, and such an evaluation set needs to be developed.
DOI: 10.1007/s10479-019-03337-5
发表时间: 2020-11-01
影响因子: 4.8
作者:
Lampiris, Georgios;Karelakis, Christos;Loizou, Efstratios
通讯作者: Loizou, Efstratios
DOI: 10.1044/2019_ajslp-cac48-18-0220
发表时间: 2020-02-01
影响因子: 2.6
作者:
Lee, Jaime B.;Azios, Jamie H.
通讯作者: Azios, Jamie H.