Veracity of Big Data: Challenges of Cross-Modal Truth Discovery

Veracity of Big Data: Challenges of Cross-Modal Truth Discovery
复制标题

大数据的准确性:跨模式真相发现的挑战

DOI:
10.1145/2935753
复制
发表时间:
2016
期刊:
ACM J. Data Inf. Qual.
影响因子:
--
通讯作者:
M. Ba
M. Ba
中科院分区:
--
文献类型:
--
作者:
Laure Berti;M. Ba

文献摘要

被引文献

相似文献

随着在线用户生成的内容呈指数级增长,对网络和社交媒体数据的依赖也在增加。从网络中发现真相具有重要的现实意义,因为网络谣言和错误信息会对我们的社会和日常生活产生巨大影响。一个根本的困难是,数据可能是有偏见的,嘈杂的,过时的,不正确的,误导性的,因此不可靠。从多个来源收集数据加剧了这一问题,必须估计数据的准确性。除了新兴的计算新闻领域和在线事实核查的成功(例如,FactCheck 1 ClaimBuster 2),真相发现是人工智能,数据库和复杂系统中许多研究社区研究的一个长期存在且具有挑战性的问题,并以各种名称进行研究:事实检查,数据和知识融合,信息可信度,可信度和信息确证(请参阅Berti-Equille和Borge-Holthoefer [2015]的调查)。其最终目标是在没有或只有很少先验知识的情况下,预测多个信息源声明的断言集合的真值标签,并推断源的可靠性。以前的一个主要工作旨在迭代计算和更新源的可信度作为其声明中的信念函数,然后每个声明的信念得分作为其源的可信度的函数[Yin et al. 2008]。然后,更复杂的概率模型包含了来源可信度和声明信念之外的各个方面,例如来源之间的依赖性,声明的相关性[Pochampally et al. 2014]以及进化真理的概念。近期缴费进一步放松
As online user-generated content grows exponentially, the reliance on web and social media data is increasing. Truth discovery from the web has significant practical importance as online rumors and misinformation can have tremendous impacts on our society and everyday life. One of the fundamental difficulties is that data can be biased, noisy, outdated, incorrect, misleading, and thus unreliable. Conflicting data from multiple sources amplifies this problem, and the veracity of data has to be estimated. Beyond the emerging field of computational journalism and the success of online fact-checkers (e.g., FactCheck1 ClaimBuster2), truth discovery is a long-standing and challenging problem studied by many research communities in artificial intelligence, databases, and complex systems, and under various names: fact-checking, data and knowledge fusion, information trustworthiness, credibility, and information corroboration (see Berti-Equille and Borge-Holthoefer [2015] for a survey). The ultimate goal is to predict the truth label of a set of assertions claimed by multiple information sources and to infer sources’ reliability with no or little prior knowledge. One major line of previous work aimed at iteratively computing and updating the source’s trustworthiness as a belief function in its claims, and then the belief score of each claim as a function of its sources’ trustworthiness [Yin et al. 2008]. More complex probabilistic models have then incorporated various aspects beyond source trustworthiness and claim belief, such as the dependence between sources, the correlation of claims [Pochampally et al. 2014], and the notion of evolving truth. Recent contributions have further relaxed