Beta Probabilistic Databases: A Scalable Approach to Belief Updating and Parameter Learning

Beta Probabilistic Databases: A Scalable Approach to Belief Updating and Parameter Learning
复制标题

Beta 概率数据库:一种可扩展的置信更新和参数学习方法

DOI:
10.1145/3035918.3064026
复制
发表时间:
2017
期刊:
Proceedings of the 2017 ACM International Conference on Management of Data
影响因子:
--
通讯作者:
Gatterbauer, Wolfgang
Gatterbauer, Wolfgang
中科院分区:
--
文献类型:
--
作者:
Meneghetti, Niccolo';Kennedy, Oliver;Gatterbauer, Wolfgang

文献摘要

参考文献

被引文献

相似文献

元组独立概率数据库(TI-PDB)通过用概率参数注释每个元组来处理不确定性;当用户提交查询时,数据库导出每个输出元组的边际概率,假设输入元组在统计上是独立的。虽然TI-PDB中的查询处理已被广泛研究,但有限的研究一直致力于更新或从查询结果的观察中获取参数的问题。解决这一问题是本文的主要重点。我们介绍了Beta概率数据库(B-PDB),TI-PDB的推广,旨在支持(i)信念更新和(ii)参数学习的原则和可扩展的方式。B-PDB的关键思想是将每个参数视为潜在的Beta分布随机变量。我们展示了这种简单的权宜之计如何以原则性的方式实现信念更新和参数学习,而不会对常规查询处理造成任何负担。我们使用这个模型提供了以下关键贡献:(i)我们展示了如何可扩展地计算新证据参数的后验密度;(ii)我们研究了执行贝叶斯信念更新的复杂性,为易处理的查询类设计了有效的算法;(iii)我们提出了一个软EM算法来计算参数的最大似然估计;(iv)我们展示了如何将所提出的算法嵌入到标准的关系引擎中;(v)我们用大量的实验结果支持我们的结论。
Tuple-independent probabilistic databases (TI-PDBs) handle uncertainty by annotating each tuple with a probability parameter; when the user submits a query, the database derives the marginal probabilities of each output-tuple, assuming input-tuples are statistically independent. While query processing in TI-PDBs has been studied extensively, limited research has been dedicated to the problems ofupdating or deriving the parameters from observations of query results. Addressing this problem is the main focus of this paper. We introduceBeta Probabilistic Databases(B-PDBs), a generalization of TI-PDBs designed to support both (i)belief updatingand (ii)parameter learningin a principled and scalable way. The key idea of B-PDBs is to treat each parameter as a latent, Beta-distributed random variable. We show how this simple expedient enables both belief updating and parameter learning in a principled way, without imposing any burden on regular query processing. We use this model to provide the following key contributions: (i) we show how to scalably compute the posterior densities of the parameters given new evidence; (ii) we study the complexity of performing Bayesian belief updates, devising efficient algorithms for tractable classes of queries; (iii) we propose a soft-EM algorithm for computing maximum-likelihood estimates of the parameters; (iv) we show how to embed the proposed algorithms into a standard relational engine; (v) we support our conclusions with extensive experimental results.
DOI: 10.1007/s00109-021-02166-z
发表时间: 2022-03
期刊: Journal of molecular medicine (Berlin, Germany)
影响因子: --
作者:
Im NR;Kim B;Jung KY;Baek SK
通讯作者: Baek SK
使用概率数据库进行近似提升推理
DOI: 10.14778/2735479.2735494
发表时间: 2014
期刊: Proc. VLDB Endow.
影响因子: --
作者:
Wolfgang Gatterbauer;Dan Suciu
通讯作者: Dan Suciu
SPROUT2:用于不确定网络数据的平方查询引擎
DOI: 10.1145/1989323.1989481
发表时间: 2011
期刊: ACM Trans. Database Syst.
影响因子: --
作者:
Robert Fink;A. Hogue;Dan Olteanu;Swaroop Rath
通讯作者: Swaroop Rath
Trio 数据不确定性和沿袭系统
DOI: 10.1007/978-0-387-09690-2_5
发表时间: 2009
期刊: Machine Learning
影响因子: 7.5
作者:
C. Aggarwal
通讯作者: C. Aggarwal
通过youtopia系统查询进行协调
DOI: --
发表时间: 2011
期刊: ACM SIGMOD Conference
影响因子: --
作者:
Nitin Gupta;Lucja Kot;G. Bender;Sudip Roy;J. Gehrke;Christoph E. Koch
通讯作者: Christoph E. Koch