UTDRM: unsupervised method for training debunked-narrative retrieval models
UTDRM: unsupervised method for training debunked-narrative retrieval models
复制标题
DOI:
10.1140/epjds/s13688-023-00437-y
复制
发表时间:
2023-12
期刊:
影响因子:
3.6
通讯作者:
Iknoor Singh;Carolina Scarton;Kalina Bontcheva
中科院分区:
文献类型:
--
作者:
Iknoor Singh;Carolina Scarton;Kalina Bontcheva
A key task in the fact-checking workflow is to establish whether the claim under investigation has already been debunked or fact-checked before. This is essentially a retrieval task where a misinformation claim is used as a query to retrieve from a corpus of debunks. Prior debunk retrieval methods have typically been trained on annotated pairs of misinformation claims and debunks. The novelty of this paper is an Unsupervised Method for Training Debunked-Narrative Retrieval Models (UTDRM) in a zero-shot setting, eliminating the need for human-annotated pairs. This approach leverages fact-checking articles for the generation of synthetic claims and employs a neural retrieval model for training. Our experiments show thatUTDRMtends to match or exceed the performance of state-of-the-art methods on seven datasets, which demonstrates its effectiveness and broad applicability. The paper also analyses the impact of various factors onUTDRM’s performance, such as the quantity of fact-checking articles utilised, the number of synthetically generated claims employed, the proposedentity inoculationmethod, and the usage of large language models for retrieval.