UDON: Unsupervised Data SelectiON for Biomedical Entity Recognition

UDON: Unsupervised Data SelectiON for Biomedical Entity Recognition
复制标题

UDON:用于生物医学实体识别的无监督数据选择

DOI:
10.1145/3507524.3507525
复制
发表时间:
2021
期刊:
Proceedings of 4th International Conference on Computing and Big Data (ICCBD)
影响因子:
--
通讯作者:
Shibuya Tetsuo
Shibuya Tetsuo
中科院分区:
--
文献类型:
--
作者:
Akdemir Arda;Shibuya Tetsuo

文献摘要

相似文献

高质量的训练数据集对于构建成功的基于机器学习 (ML) 的 NLP 系统至关重要。然而,这些数据集并不总是在生物医学领域等资源匮乏的环境中可用。在这里,选择相关训练数据与选择 ML 模型同样重要。在本研究中,我们提出 UDON:使用特定领域的预训练语言模型 (LM) 进行生物医学实体识别的无监督数据选择。我们首先证明,预训练的 LM 在没有任何监督的情况下成功地隐式学习数据集之间的差异,然后使用这些模型来选择相关的数据实例。接下来,我们使用四个 LM 和三种选择方法在七个生物医学数据集和一个新闻领域数据集上评估所提出的实体识别方法。我们的结果表明,使用预训练的特定领域 LM 进行数据选择优于所有其他方法。最后,我们使用域分类作为在域内数据集上预训练神经网络的辅助任务,并表明这会产生进一步的改进。
High-quality training datasets are critical for building successful Machine Learning (ML) based NLP systems. However, these datasets are not always available in low-resource contexts such as the biomedical domain. Here, selecting relevant training data is as important as the choice of the ML model. In this study we propose UDON: Unsupervised Data selectiON for biomedical entity recognition using domain-specific pretrained Language Models (LMs). We first show that pretrained LMs succeed at implicitly learning the differences between datasets without any supervision, and then use these models to select relevant data instances. Next, we evaluate the proposed methods for entity recognition on seven biomedical datasets and one news domain dataset using four LMs and three selection methods. Our results show that using pretrained domain-specific LMs for data selection outperforms all other approaches. Finally, we use domain classification as an auxiliary task for pretraining the neural network on the in-domain dataset and show this yields further improvements.