SagDRE: Sequence-Aware Graph-Based Document-Level Relation Extraction with Adaptive Margin Loss

SagDRE: Sequence-Aware Graph-Based Document-Level Relation Extraction with Adaptive Margin Loss
复制标题

DOI:
10.1145/3534678.3539304
复制
发表时间:
2022-08
期刊:
Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
Ying Wei;Qi Li
Ying Wei;Qi Li
中科院分区:
其他
文献类型:
--
作者:
Ying Wei;Qi Li

文献摘要

相似文献

关系抽取是许多自然语言处理应用的重要任务。文档级关系抽取任务的目的是抽取文档内部的关系,这对逆向工程任务提出了许多挑战,因为它需要跨句子推理和处理同一文档中表达的多个关系。现有最先进的文档级RE模型使用图形结构来更好地连接远程关联。在这项工作中,我们提出了SagDRE模型,该模型进一步考虑并捕获了文本中的原始序列信息。该模型通过学习语句级的方向边来捕捉文档中的信息流,并使用令牌级的顺序信息来编码从一个实体到另一个实体的最短路径。此外,我们还提出了一种自适应余量损失来解决文档级逆向工程任务的长尾多标签问题,其中一个实体对可以在一个文档中表示多个关系,并且存在一些流行的关系。损失函数旨在鼓励积极和消极类别之间的分离。在不同领域的数据集上的实验结果证明了该方法的有效性。
Relation extraction (RE) is an important task for many natural language processing applications. Document-level relation extraction task aims to extract the relations within a document and poses many challenges to the RE tasks as it requires reasoning across sentences and handling multiple relations expressed in the same document. Existing state-of-the-art document-level RE models use the graph structure to better connect long-distance correlations. In this work, we propose SagDRE model, which further considers and captures the original sequential information from the text. The proposed model learns sentence-level directional edges to capture the information flow in the document and uses the token-level sequential information to encode the shortest paths from one entity to the other. In addition, we propose an adaptive margin loss to address the long-tailed multi-label problem of document-level RE tasks, where multiple relations can be expressed in a document for an entity pair and there are a few popular relations. The loss function aims to encourage separations between positive and negative classes. The experimental results on datasets from various domains demonstrate the effectiveness of the proposed methods.