BioBERT: a pre-trained biomedical language representation model for biomedical text mining.

BioBERT: a pre-trained biomedical language representation model for biomedical text mining.
复制标题

DOI:
10.1093/bioinformatics/btz682
复制
发表时间:
2020-02-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Kang J
Kang J
中科院分区:
其他
文献类型:
--
作者:
Lee J;Yoon W;Kim S;Kim D;Kim S;So CH;Kang J

文献摘要

参考文献

被引文献

相似文献

随着生物医学文档数量的快速增长,生物医学文本挖掘变得越来越重要。随着自然语言处理(NLP)的进步,从生物医学文献中提取有价值的信息已经受到研究人员的欢迎,深度学习促进了有效的生物医学文本挖掘模型的发展。然而,直接将NLP的进步应用于生物医学文本挖掘往往会产生不满意的结果,这是由于单词分布从一般领域语料库转移到生物医学语料库。在这篇文章中,我们研究了最近引入的预训练语言模型BERT如何适用于生物医学语料库。我们介绍了BioBERT(Bidirectional Encoder Representations from Transformers for Biomedical Text Mining),这是一个在大规模生物医学语料库上预训练的特定领域语言表示模型。在任务之间具有几乎相同的架构,当在生物医学语料库上进行预训练时,BioBERT在各种生物医学文本挖掘任务中的表现大大优于BERT和以前的最先进模型。虽然BERT获得的性能与以前的最先进的模型相当,但BioBERT在以下三个代表性的生物医学文本挖掘任务上显着优于它们:生物医学命名实体识别(0.62%F1分数提高),生物医学关系提取(2.80%F1分数提高)和生物医学问题回答(12.24%MRR提高)。我们的分析结果表明,生物医学语料库上的预训练BERT有助于它理解复杂的生物医学文本。我们在https://github.com/naver/biobert-pretrained上免费提供BioBERT的预训练权重,并在https://github.com/dmis-lab/biobert上提供微调BioBERT的源代码。
Biomedical text mining is becoming increasingly important as the number of biomedical documents rapidly grows. With the progress in natural language processing (NLP), extracting valuable information from biomedical literature has gained popularity among researchers, and deep learning has boosted the development of effective biomedical text mining models. However, directly applying the advancements in NLP to biomedical text mining often yields unsatisfactory results due to a word distribution shift from general domain corpora to biomedical corpora. In this article, we investigate how the recently introduced pre-trained language model BERT can be adapted for biomedical corpora. We introduce BioBERT (Bidirectional Encoder Representations from Transformers for Biomedical Text Mining), which is a domain-specific language representation model pre-trained on large-scale biomedical corpora. With almost the same architecture across tasks, BioBERT largely outperforms BERT and previous state-of-the-art models in a variety of biomedical text mining tasks when pre-trained on biomedical corpora. While BERT obtains performance comparable to that of previous state-of-the-art models, BioBERT significantly outperforms them on the following three representative biomedical text mining tasks: biomedical named entity recognition (0.62% F1 score improvement), biomedical relation extraction (2.80% F1 score improvement) and biomedical question answering (12.24% MRR improvement). Our analysis results show that pre-training BERT on biomedical corpora helps it to understand complex biomedical texts. We make the pre-trained weights of BioBERT freely available at https://github.com/naver/biobert-pretrained, and the source code for fine-tuning BioBERT available at https://github.com/dmis-lab/biobert.
DOI: 10.1371/journal.pone.0065390
发表时间: 2013
期刊: PloS one
影响因子: 3.7
作者:
Pafilis E;Frankild SP;Fanini L;Faulwetter S;Pavloudi C;Vasileiadou A;Arvanitidis C;Jensen LJ
通讯作者: Jensen LJ
基于注意力的 BiLSTM-CRF 文档级化学命名实体识别方法
DOI: 10.1093/bioinformatics/btx761
发表时间: 2018-04-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Luo, Ling;Yang, Zhihao;Wang, Jian
通讯作者: Wang, Jian
DOI: 10.1186/s12859-015-0472-9
发表时间: 2015-02-21
期刊: BMC bioinformatics
影响因子: 3
作者:
Bravo À;Piñero J;Queralt-Rosinach N;Rautschka M;Furlong LI
通讯作者: Furlong LI
DOI: 10.1093/bioinformatics/btx228
发表时间: 2017-07-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Habibi M;Weber L;Neves M;Wiegandt DL;Leser U
通讯作者: Leser U
DOI: 10.1016/j.jbi.2013.12.006
发表时间: 2014-02
影响因子: 4.5
作者:
Dogan, Rezarta Islamaj;Leaman, Robert;Lu, Zhiyong
通讯作者: Lu, Zhiyong