Unsupervised Partial Sentence Matching for Cited Text Identification

Unsupervised Partial Sentence Matching for Cited Text Identification
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Kathryn Ricci;Haw-Shiuan Chang;Purujit Goyal;A. McCallum
Kathryn Ricci;Haw-Shiuan Chang;Purujit Goyal;A. McCallum
中科院分区:
其他
文献类型:
--
作者:
Kathryn Ricci;Haw-Shiuan Chang;Purujit Goyal;A. McCallum

文献摘要

相似文献

在研究论文正文中给出一条引文,被引文本识别的目的是在被引论文中找到与被引句子最相关的句子。这项任务从根本上说是一项句子匹配任务,其中亲和力通常通过句子嵌入之间的余弦相似度来评估。然而,(A)单个嵌入可能不能很好地表示句子,因为它们包含多个不同的语义方面,以及(B)好的匹配可能不需要在所有方面都有很强的匹配。为了克服这些局限性,我们提出了一种简单有效的无监督引用文本识别方法,该方法采用非对称相似性度量,允许两个句子中多个方面的部分匹配。在CL-SciSumm数据集上,我们发现我们的方法优于基线对称方法,令人惊讶的是,我们的方法也优于提交给CL-SciSumm共享任务1a的过去版本的所有监督和非监督系统。
Given a citation in the body of a research paper, cited text identification aims to find the sentences in the cited paper that are most relevant to the citing sentence. The task is fundamentally one of sentence matching, where affinity is often assessed by a cosine similarity between sentence embeddings. However, (a) sentences may not be well-represented by a single embedding because they contain multiple distinct semantic aspects, and (b) good matches may not require a strong match in all aspects. To overcome these limitations, we propose a simple and efficient unsupervised method for cited text identification that adapts an asymmetric similarity measure to allow partial matches of multiple aspects in both sentences. On the CL-SciSumm dataset we find that our method outperforms a baseline symmetric approach, and, surprisingly, also outperforms all supervised and unsupervised systems submitted to past editions of CL-SciSumm Shared Task 1a.