SLiMFast: Guaranteed Results for Data Fusion and Source Reliability

SLiMFast: Guaranteed Results for Data Fusion and Source Reliability
复制标题

DOI:
10.1145/3035918.3035951
复制
发表时间:
2015-12
期刊:
Proceedings of the 2017 ACM International Conference on Management of Data
影响因子:
--
通讯作者:
Theodoros Rekatsinas;Manas R. Joglekar;H. Garcia-Molina;Aditya G. Parameswaran;Christopher Ré
Theodoros Rekatsinas;Manas R. Joglekar;H. Garcia-Molina;Aditya G. Parameswaran;Christopher Ré
中科院分区:
其他
文献类型:
--
作者:
Theodoros Rekatsinas;Manas R. Joglekar;H. Garcia-Molina;Aditya G. Parameswaran;Christopher Ré

文献摘要

被引文献

相似文献

我们专注于数据融合,即通过估计数据源精度将数据源中的冲突数据统一为单个表示的问题。我们提出了SLiMFast框架,该框架将数据融合表达为判别概率模型上的统计学习问题,在许多情况下对应于逻辑回归。与之前使用复杂生成模型的方法相比,判别模型对数据源的分布假设较少,并允许我们获得严格的理论保证。此外,我们还展示了SLiMFast如何将领域知识整合到数据融合中,从而比最先进的基线提高了高达50%的准确性。基于我们的理论结果,我们设计了一个优化器,避免了用户手动选择学习SLiMFast参数的算法。我们在多个真实数据集上验证了我们的优化器,并表明它可以准确地预测产生最佳数据融合结果的学习算法。
We focus on data fusion, i.e., the problem of unifying conflicting data from data sources into a single representation by estimating the source accuracies. We propose SLiMFast, a framework that expresses data fusion as a statistical learning problem over discriminative probabilistic models, which in many cases correspond to logistic regression. In contrast to previous approaches that use complex generative models, discriminative models make fewer distributional assumptions over data sources and allow us to obtain rigorous theoretical guarantees. Furthermore, we show how SLiMFast enables incorporating domain knowledge into data fusion, yielding accuracy improvements of up to 50% over state-of-the-art baselines. Building upon our theoretical results, we design an optimizer that obviates the need for users to manually select an algorithm for learning SLiMFast's parameters. We validate our optimizer on multiple real-world datasets and show that it can accurately predict the learning algorithm that yields the best data fusion results.