A Richly Annotated Corpus for Different Tasks in Automated Fact-Checking

A Richly Annotated Corpus for Different Tasks in Automated Fact-Checking
复制标题

DOI:
10.18653/v1/k19-1046
复制
发表时间:
2019-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Andreas Hanselowski;Christian Stab;Claudia Schulz;Zile Li;Iryna Gurevych
Andreas Hanselowski;Christian Stab;Claudia Schulz;Zile Li;Iryna Gurevych
中科院分区:
其他
文献类型:
--
作者:
Andreas Hanselowski;Christian Stab;Claudia Schulz;Zile Li;Iryna Gurevych

文献摘要

被引文献

相似文献

基于机器学习的自动事实检查是一种有前途的方法,可以识别在网络上分发的虚假信息。为了实现令人满意的性能,机器学习方法需要大型语料库,并在事实检查过程中对不同任务进行可靠的注释。在分析了现有的事实检验语料库之后,我们发现它们都不符合这些标准。它们的尺寸太小,不提供详细的注释,或者仅限于单个域。在这一差距的推动下,我们提出了一个新的大小的混合域语料库,并具有良好质量的核心事实检查任务的注释:文件检索,证据提取,立场检测和索赔验证。为了帮助未来的语料库构建,我们描述了我们的语料库创建和注释方法,并证明它导致了大量的通知者一致。作为未来研究的基础,我们在语料库上进行了许多模型体系结构,这些模型体系结构在类似的问题设置中达到了高性能。最后,为了支持未来模型的开发,我们为每个任务提供了详细的错误分析。我们的结果表明,由我们的数据定义的现实,多域设置为现有模型带来了新的挑战,为未来系统提供了可观的改进机会。
Automated fact-checking based on machine learning is a promising approach to identify false information distributed on the web. In order to achieve satisfactory performance, machine learning methods require a large corpus with reliable annotations for the different tasks in the fact-checking process. Having analyzed existing fact-checking corpora, we found that none of them meets these criteria in full. They are either too small in size, do not provide detailed annotations, or are limited to a single domain. Motivated by this gap, we present a new substantially sized mixed-domain corpus with annotations of good quality for the core fact-checking tasks: document retrieval, evidence extraction, stance detection, and claim validation. To aid future corpus construction, we describe our methodology for corpus creation and annotation, and demonstrate that it results in substantial inter-annotator agreement. As baselines for future research, we perform experiments on our corpus with a number of model architectures that reach high performance in similar problem settings. Finally, to support the development of future models, we provide a detailed error analysis for each of the tasks. Our results show that the realistic, multi-domain setting defined by our data poses new challenges for the existing models, providing opportunities for considerable improvement by future systems.