Multi-Vector Models with Textual Guidance for Fine-Grained Scientific Document Similarity

Multi-Vector Models with Textual Guidance for Fine-Grained Scientific Document Similarity
复制标题

DOI:
10.18653/v1/2022.naacl-main.331
复制
发表时间:
2021-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Sheshera Mysore;Arman Cohan;Tom Hope
Sheshera Mysore;Arman Cohan;Tom Hope
中科院分区:
其他
文献类型:
--
作者:
Sheshera Mysore;Arman Cohan;Tom Hope

文献摘要

相似文献

提出了一种基于文本细粒度特征匹配的科技文档相似度模型。为了训练我们的模型,我们利用了一个自然发生的监督来源:论文全文中的句子,这些句子一起引用了多篇论文(共同引用)。这样的共同引用不仅反映了密切的论文相关性,而且还提供了共同引用的论文是如何相关的文本描述。这种新颖的文本监督形式用于学习跨论文匹配方面。我们开发了多向量表示,其中向量对应于文档的文档级方面,并提出了两种方面匹配方法:(1)仅匹配单个方面的快速方法,以及(2)使用最优传输机制进行稀疏多个匹配的方法,该机制计算方面之间的地球移动器距离。我们的方法提高了四个数据集的文档相似性任务的性能。此外,我们的快速单匹配方法取得了有竞争力的结果,为将细粒度相似性应用于大型科学语料库铺平了道路。
We present a new scientific document similarity model based on matching fine-grained aspects of texts. To train our model, we exploit a naturally-occurring source of supervision: sentences in the full-text of papers that cite multiple papers together (co-citations). Such co-citations not only reflect close paper relatedness, but also provide textual descriptions of how the co-cited papers are related. This novel form of textual supervision is used for learning to match aspects across papers. We develop multi-vector representations where vectors correspond to sentence-level aspects of documents, and present two methods for aspect matching: (1) A fast method that only matches single aspects, and (2) a method that makes sparse multiple matches with an Optimal Transport mechanism that computes an Earth Mover’s Distance between aspects. Our approach improves performance on document similarity tasks in four datasets. Further, our fast single-match method achieves competitive results, paving the way for applying fine-grained similarity to large scientific corpora.