We are not ready yet: limitations of transfer learning for Disease Named Entity Recognition

We are not ready yet: limitations of transfer learning for Disease Named Entity Recognition
复制标题

我们还没有准备好:疾病命名实体识别的迁移学习的局限性

DOI:
10.1101/2021.07.11.451939
复制
发表时间:
2021
期刊:
bioRxiv
影响因子:
--
通讯作者:
J. Fluck
J. Fluck
中科院分区:
--
文献类型:
--
作者:
Lisa Langnickel;J. Fluck

文献摘要

被引文献

相似文献

在生物医学自然语言处理领域进行了大量的研究。自从基于迁移学习的方法取得突破以来,BERT模型被用于各种生物医学和临床应用。对于可用的数据集,这些模型显示了出色的结果-部分超过了注释者之间的协议。然而,与现有测试数据的结果相比,应用于COVID-19预印本的生物医学命名实体识别显示出性能下降。问题是,经过良好训练的模型如何能够对全新的数据进行预测,即进行概括。基于疾病命名实体识别的例子,我们研究了不同的基于机器学习的方法(即迁移学习)的鲁棒性,并表明当前最先进的方法对于给定的训练和相应的测试集都能很好地工作,但在应用于新数据时却明显缺乏泛化能力。因此,我们认为,有必要为训练和测试更大的注释数据集。
Intense research has been done in the area of biomedical natural language processing. Since the breakthrough of transfer learning-based methods, BERT models are used in a variety of biomedical and clinical applications. For the available data sets, these models show excellent results – partly exceeding the inter-annotator agreements. However, biomedical named entity recognition applied on COVID-19 preprints shows a performance drop compared to the results on available test data. The question arises how well trained models are able to predict on completely new data, i.e. to generalize. Based on the example of disease named entity recognition, we investigate the robustness of different machine learning-based methods – thereof transfer learning – and show that current state-of-the-art methods work well for a given training and the corresponding test set but experience a significant lack of generalization when applying to new data. We therefore argue that there is a need for larger annotated data sets for training and testing.