BioREx: Improving biomedical relation extraction by leveraging heterogeneous datasets

BioREx: Improving biomedical relation extraction by leveraging heterogeneous datasets
复制标题

DOI:
10.1016/j.jbi.2023.104487
复制
发表时间:
2023-09-07
影响因子:
4.5
通讯作者:
Lu,Zhiyong
Lu,Zhiyong
中科院分区:
医学3区
文献类型:
--
作者:
Lai,Po-Ting;Wei,Chih-Hsuan;Lu,Zhiyong

文献摘要

相似文献

生物医学关系抽取是从自由文本中自动识别和刻画生物医学概念之间关系的任务。逆向工程是生物医学自然语言处理(NLP)研究的中心任务,在基于文献的发现和知识图构建等后续应用中起着至关重要的作用。最先进的方法主要用于在单个RE数据集上训练机器学习模型,例如蛋白质-蛋白质相互作用和化学诱导的疾病关系。然而,手动数据集注释非常昂贵且耗时,因为它需要领域知识。现有的反求工程数据集通常是特定领域的或较小的,这限制了通用和高性能反求工程模型的发展。在这项工作中,我们提出了一种新的框架,用于系统地解决单个数据集的数据异构性,并将它们合并为一个大型数据集。基于该框架和数据集,我们报告了BioREx,这是一种以数据为中心的关系提取方法。我们的评估表明,BioREx的性能明显高于在单个数据集上训练的基准系统,在最近发布的BioRED语料库上,新的SOTA在F-1测试中从74.4%到79.6%。我们进一步证明,对于五个不同的RE任务,组合数据集可以提高性能。此外,我们表明,平均而言,BioREx比目前表现最好的方法,如迁移学习和多任务学习具有更好的优势。最后,我们在两个以前在训练数据中未见的独立RE任务中展示了BioREx的健壮性和泛化能力:药物-药物N-ary组合和文档级基因-疾病RE。集成数据集和优化方法已打包为https://github.com/ncbi/BioREx.提供的独立工具
Biomedical relation extraction (RE) is the task of automatically identifying and characterizing relations between biomedical concepts from free text. RE is a central task in biomedical natural language processing (NLP) research and plays a critical role in many downstream applications, such as literature-based discovery and knowledge graph construction. State-of-the-art methods were used primarily to train machine learning models on individual RE datasets, such as protein–protein interaction and chemical-induced disease relation. Manual dataset annotation, however, is highly expensive and time-consuming, as it requires domain knowledge. Existing RE datasets are usually domain-specific or small, which limits the development of generalized and high-performing RE models. In this work, we present a novel framework for systematically addressing the data heterogeneity of individual datasets and combining them into a large dataset. Based on the framework and dataset, we report on BioREx, a data-centric approach for extracting relations. Our evaluation shows that BioREx achieves significantly higher performance than the benchmark system trained on the individual dataset, setting a new SOTA from 74.4% to 79.6% in F-1 measure on the recently released BioRED corpus. We further demonstrate that the combined dataset can improve performance for five different RE tasks. In addition, we show that on average BioREx compares favorably to current best-performing methods such as transfer learning and multi-task learning. Finally, we demonstrate BioREx’s robustness and generalizability in two independent RE tasks not previously seen in training data: drug-drug N-ary combination and document-level gene-disease RE. The integrated dataset and optimized method have been packaged as a stand-alone tool available at https://github.com/ncbi/BioREx.