MultiFC: A Real-World Multi-Domain Dataset for Evidence-Based Fact Checking of Claims

MultiFC: A Real-World Multi-Domain Dataset for Evidence-Based Fact Checking of Claims
复制标题

DOI:
10.18653/v1/d19-1475
复制
发表时间:
2019-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Isabelle Augenstein;C. Lioma;Dongsheng Wang-;Lucas Chaves Lima;Casper Hansen;Christian Hansen;J. Simonsen
Isabelle Augenstein;C. Lioma;Dongsheng Wang-;Lucas Chaves Lima;Casper Hansen;Christian Hansen;J. Simonsen
中科院分区:
其他
文献类型:
--
作者:
Isabelle Augenstein;C. Lioma;Dongsheng Wang-;Lucas Chaves Lima;Casper Hansen;Christian Hansen;J. Simonsen

文献摘要

被引文献

相似文献

我们贡献了自然发生的事实索赔的最大公开数据集,用于自动索赔验证。它收集自26个英文事实核查网站,与文本来源和丰富的元数据配对,并由人类专家记者标记其准确性。我们对数据集进行了深入分析,突出了特征和挑战。此外,我们提出了自动准确性预测的结果,既有既定的基线,也有一种新的方法,用于证据页面的联合排名和预测准确性,优于所有基线。通过对证据进行编码和对元数据进行建模,实现了显著的性能提升。我们表现最好的模型达到了49.2%的宏观F1,这表明这是一个具有挑战性的索赔准确性预测测试平台。
We contribute the largest publicly available dataset of naturally occurring factual claims for the purpose of automatic claim verification. It is collected from 26 fact checking websites in English, paired with textual sources and rich metadata, and labelled for veracity by human expert journalists. We present an in-depth analysis of the dataset, highlighting characteristics and challenges. Further, we present results for automatic veracity prediction, both with established baselines and with a novel method for joint ranking of evidence pages and predicting veracity that outperforms all baselines. Significant performance increases are achieved by encoding evidence, and by modelling metadata. Our best-performing model achieves a Macro F1 of 49.2%, showing that this is a challenging testbed for claim veracity prediction.