Probabilistic data exchange

Probabilistic data exchange
复制标题

概率数据交换

DOI:
10.1145/1804669.1804681
复制
发表时间:
2011
期刊:
J. ACM
影响因子:
--
通讯作者:
Phokion G. Kolaitis
Phokion G. Kolaitis
中科院分区:
--
文献类型:
--
作者:
Ronald Fagin;B. Kimelfeld;Phokion G. Kolaitis

文献摘要

被引文献

相似文献

这里报告的工作为存在概率数据的数据交换奠定了基础。这需要重新思考传统数据交换的基本概念,如解决方案、通用解决方案和目标查询的特定答案。我们开发了一个在概率数据库上进行数据交换的框架,并对其一致性和鲁棒性进行了论证。该框架适用于任意模式映射,以及源和目标实例上的有限或可数无限概率空间。在建立了该框架并阐述了关键概念之后,我们研究了该框架在一个具体的实际环境中的应用,在这个环境中,概率数据库通过在随机布尔变量上制定的注释来紧凑地编码。在这种情况下,我们研究了解决方案和通用解决方案的存在性测试,实现这些解决方案,以及在精确意义和近似意义上评估目标查询(对于合取查询的并集)的问题。对于每个问题,我们在各种依赖类中基于注释的属性执行复杂性分析。最后,我们证明了框架和结果很容易和完全泛化,不仅允许数据,而且允许模式映射本身是概率的。
The work reported here lays the foundations of data exchange in the presence of probabilistic data. This requires rethinking the very basic concepts of traditional data exchange, such as solution, universal solution, and the certain answers of target queries. We develop a framework for data exchange over probabilistic databases, and make a case for its coherence and robustness. This framework applies to arbitrary schema mappings, and finite or countably infinite probability spaces on the source and target instances. After establishing this framework and formulating the key concepts, we study the application of the framework to a concrete and practical setting where probabilistic databases are compactly encoded by means of annotations formulated over random Boolean variables. In this setting, we study the problems of testing for the existence of solutions and universal solutions, materializing such solutions, and evaluating target queries (for unions of conjunctive queries) in both the exact sense and the approximate sense. For each of the problems, we carry out a complexity analysis based on properties of the annotation, in various classes of dependencies. Finally, we show that the framework and results easily and completely generalize to allow not only the data, but also the schema mapping itself to be probabilistic.