Improving Supervised Drug-Protein Relation Extraction with Distantly Supervised Models

Improving Supervised Drug-Protein Relation Extraction with Distantly Supervised Models
复制标题

DOI:
10.18653/v1/2022.bionlp-1.16
复制
发表时间:
2022
期刊:
Journal of Physics: Conference Series
影响因子:
--
通讯作者:
Naoki Iinuma;Makoto Miwa;Yutaka Sasaki
Naoki Iinuma;Makoto Miwa;Yutaka Sasaki
中科院分区:
其他
文献类型:
--
作者:
Naoki Iinuma;Makoto Miwa;Yutaka Sasaki

文献摘要

相似文献

本文提出了新的药物蛋白质关系提取模型,间接利用远程监督数据。具体地说,我们的模型不是将远程监督数据添加到手动注释的训练数据中,而是将远程监督模型纳入其中,这些模型是用远程监督数据训练的关系提取模型。已经提出了远程监督学习以低成本生成大量的伪训练数据。然而,由于包含错误标记的数据,仍然存在预测性能低的问题。因此,已经提出了几种方法来抑制噪声的情况下,通过利用一些手动注释的训练数据的影响。然而,它们的性能低于手动注释数据的监督学习,因为无法完全抑制的错误标记数据在训练模型时会变成噪声。为了克服这个问题,我们的方法间接利用远程监督数据与手动注释的训练数据。在BioCreative VII Track 1中的DrugProt语料库上的实验结果表明,我们提出的模型可以在不同的环境中持续改进监督模型。
This paper proposes novel drug-protein relation extraction models that indirectly utilize distant supervision data. Concretely, instead of adding distant supervision data to the manually annotated training data, our models incorporate distantly supervised models that are relation extraction models trained with distant supervision data. Distantly supervised learning has been proposed to generate a large amount of pseudo-training data at low cost. However, there is still a problem of low prediction performance due to the inclusion of mislabeled data. Therefore, several methods have been proposed to suppress the effects of noisy cases by utilizing some manually annotated training data. However, their performance is lower than that of supervised learning on manually annotated data because mislabeled data that cannot be fully suppressed becomes noise when training the model. To overcome this issue, our methods indirectly utilize distant supervision data with manually annotated training data. The experimental results on the DrugProt corpus in the BioCreative VII Track 1 showed that our proposed model can consistently improve the supervised models in different settings.