Veracity of Big Data: Challenges of Cross-Modal Truth Discovery
Veracity of Big Data: Challenges of Cross-Modal Truth Discovery
复制标题
大数据的准确性:跨模式真相发现的挑战
DOI:
10.1145/2935753
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
M. Ba
中科院分区:
文献类型:
--
作者:
Laure Berti;M. Ba
As online user-generated content grows exponentially, the reliance on web and social media data is increasing. Truth discovery from the web has significant practical importance as online rumors and misinformation can have tremendous impacts on our society and everyday life. One of the fundamental difficulties is that data can be biased, noisy, outdated, incorrect, misleading, and thus unreliable. Conflicting data from multiple sources amplifies this problem, and the veracity of data has to be estimated. Beyond the emerging field of computational journalism and the success of online fact-checkers (e.g., FactCheck1 ClaimBuster2), truth discovery is a long-standing and challenging problem studied by many research communities in artificial intelligence, databases, and complex systems, and under various names: fact-checking, data and knowledge fusion, information trustworthiness, credibility, and information corroboration (see Berti-Equille and Borge-Holthoefer [2015] for a survey). The ultimate goal is to predict the truth label of a set of assertions claimed by multiple information sources and to infer sources’ reliability with no or little prior knowledge. One major line of previous work aimed at iteratively computing and updating the source’s trustworthiness as a belief function in its claims, and then the belief score of each claim as a function of its sources’ trustworthiness [Yin et al. 2008]. More complex probabilistic models have then incorporated various aspects beyond source trustworthiness and claim belief, such as the dependence between sources, the correlation of claims [Pochampally et al. 2014], and the notion of evolving truth. Recent contributions have further relaxed