UTDRM: unsupervised method for training debunked-narrative retrieval models

UTDRM: unsupervised method for training debunked-narrative retrieval models
复制标题

DOI:
10.1140/epjds/s13688-023-00437-y
复制
发表时间:
2023-12
期刊:
影响因子:
3.6
通讯作者:
Iknoor Singh;Carolina Scarton;Kalina Bontcheva
Iknoor Singh;Carolina Scarton;Kalina Bontcheva
中科院分区:
计算机科学3区
文献类型:
--
作者:
Iknoor Singh;Carolina Scarton;Kalina Bontcheva

文献摘要

相似文献

事实核查工作流程中的一项关键任务是确定正在调查的索赔之前是否已被揭穿或经过事实核查。这本质上是一个检索任务,其中错误信息声明被用作从揭穿语料库中检索的查询。先前的揭穿检索方法通常是在带注释的错误信息声明和揭穿对上进行训练的。本文的新颖之处在于一种在零样本设置中训练揭穿叙事检索模型(UTDRM)的无监督方法,消除了对人工注释对的需要。这种方法利用事实检查文章来生成合成声明,并采用神经检索模型进行训练。我们的实验表明,UTDRM 在七个数据集上趋于匹配或超过最先进方法的性能,这证明了其有效性和广泛的适用性。本文还分析了各种因素对 UTDRM 性能的影响,例如使用的事实检查文章的数量、使用的综合生成的声明的数量、提出的实体接种方法以及用于检索的大型语言模型的使用。
A key task in the fact-checking workflow is to establish whether the claim under investigation has already been debunked or fact-checked before. This is essentially a retrieval task where a misinformation claim is used as a query to retrieve from a corpus of debunks. Prior debunk retrieval methods have typically been trained on annotated pairs of misinformation claims and debunks. The novelty of this paper is an Unsupervised Method for Training Debunked-Narrative Retrieval Models (UTDRM) in a zero-shot setting, eliminating the need for human-annotated pairs. This approach leverages fact-checking articles for the generation of synthetic claims and employs a neural retrieval model for training. Our experiments show thatUTDRMtends to match or exceed the performance of state-of-the-art methods on seven datasets, which demonstrates its effectiveness and broad applicability. The paper also analyses the impact of various factors onUTDRM’s performance, such as the quantity of fact-checking articles utilised, the number of synthetically generated claims employed, the proposedentity inoculationmethod, and the usage of large language models for retrieval.