Supporting factual statements with evidence from the web

Supporting factual statements with evidence from the web
复制标题

用网络证据支持事实陈述

DOI:
10.1145/2396761.2398415
复制
发表时间:
2012
期刊:
Proceedings of the 21st ACM international conference on Information and knowledge management
影响因子:
--
通讯作者:
Silviu Cucerzan
Silviu Cucerzan
中科院分区:
--
文献类型:
--
作者:
C. W. Leong;Silviu Cucerzan

文献摘要

被引文献

相似文献

事实核实已经成为一项重要的任务,这是因为博客、讨论组和社交网站以及汇集了许多贡献者内容的百科全书集越来越受欢迎。我们研究了从Web上自动检索事实陈述的支持证据的任务。以维基百科为起点,我们得出了一个与支持Web文档配对的大型语句语料库,我们进一步使用这些语料库作为训练和测试数据,假设对维基百科的引用代表了一些与支持相应语句最相关的Web文档。在给定事实陈述的情况下,该系统首先使用机器学习技术将其转换为一组语义短语。然后,它采用一种准随机策略来根据主题似然来选择语义短语的子集。这些语义术语用于构建用于通过Web搜索API检索Web文档的查询。最后,通过采用对检索到的文档的适宜性的附加度量来支持事实陈述,对检索到的文档进行聚合和重新排序。为了衡量检索到的证据的质量,我们通过Amazon Machine Turk进行了一项用户研究,结果表明,我们的系统能够检索到与维基百科贡献者选择的文档相当的支持Web文档。
Fact verification has become an important task due to the increased popularity of blogs, discussion groups, and social sites, as well as of encyclopedic collections that aggregate content from many contributors. We investigate the task of automatically retrieving supporting evidence from the Web for factual statements. Using Wikipedia as a starting point, we derive a large corpus of statements paired with supporting Web documents, which we employ further as training and test data under the assumption that the contributed references to Wikipedia represent some of the most relevant Web documents for supporting the corresponding statements. Given a factual statement, the proposed system first transforms it into a set of semantic terms by using machine learning techniques. It then employs a quasi-random strategy for selecting subsets of the semantic terms according to topical likelihood. These semantic terms are used to construct queries for retrieving Web documents via a Web search API. Finally, the retrieved documents are aggregated and re-ranked by employing additional measures of their suitability to support the factual statement. To gauge the quality of the retrieved evidence, we conduct a user study through Amazon Mechanical Turk, which shows that our system is capable of retrieving supporting Web documents comparable to those chosen by Wikipedia contributors.