A Novel Automated Approach to Mutation-Cancer Relation Extraction by Incorporating Heterogeneous Knowledge.

A Novel Automated Approach to Mutation-Cancer Relation Extraction by Incorporating Heterogeneous Knowledge.
复制标题

通过结合异质知识来提取突变-癌症关系的新颖自动化方法。

DOI:
10.1109/jbhi.2022.3220924
复制
发表时间:
2023
影响因子:
7.7
通讯作者:
Cao J
Cao J
中科院分区:
工程技术1区
文献类型:
--
作者:
Cao J

文献摘要

相似文献

使用文本挖掘自动提取癌症文献中出现的基因突变与癌症实体之间的关系,可以快速提供支持精准癌症医学的重要信息。然而,突变-癌症关系提取比从自由文本中提取一般关系更具挑战性,因为如果没有特定于癌症的背景知识通常是不可能的,因此该模型依赖于对复杂周围标记的更深入理解。我们提出了一种深度学习模型,可以联合提取突变及其相关癌症。背景知识来自两个不同的知识库,它们存储有关突变的不同类型的信息。考虑到知识在这两种资源中存储的方式不同,我们提出了两种不同的嵌入知识的方法,即基于句子的知识集成和属性感知的知识集成。评估表明,我们的模型优于许多基线模型,在三个公共数据集 EMU BCa、EMU PCa 和 BRONCO 上获得了 96.00%、92.57% 和 94.57% 的 F1 分数,从而说明了我们知识集成方法的有效性。辅助实验表明,尽管输入文本提供的上下文不足,但我们的模型可以利用来自知识库的更多信息文本,并将突变与其相应的癌症疾病联系起来。
Automatic extraction of relations between gene mutations and cancer entities occurring in the cancer literature using text mining can rapidly provide vital information to support precision cancer medicine. However, mutation-cancer relation extraction is more challenging than general relation extraction from free text, since it is often not possible without cancer-specific background knowledge and thus the model replies on a deeper understanding of complex surrounding tokens. We propose a deep learning model that jointly extracts mutations and their associated cancers. Background knowledge comes from two different knowledge bases which store different types of information about mutations. Given the different ways in which knowledge is stored in these two resources, we propose two separate methods for embedding knowledge, namely sentence-based knowledge integration and attribute-aware knowledge integration. The evaluation demonstrated that our model outperforms a number of baseline models and gains 96.00%, 92.57% and 94.57% F1 scores on three public datasets, EMU BCa, EMU PCa, and BRONCO, thus illustrating the effectiveness of our knowledge integration approach. The auxiliary experiments show that our models can utilize more informative text from the KBs and link the mutations to their corresponding cancer disease although the input text provides insufficient context.