Answering Frequent Probabilistic Inference Queries in Databases

Answering Frequent Probabilistic Inference Queries in Databases
复制标题

DOI:
10.1109/tkde.2010.146
复制
发表时间:
2011-04
影响因子:
8.9
通讯作者:
Shaoxu Song;Lei Chen;J. Yu
Shaoxu Song;Lei Chen;J. Yu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Shaoxu Song;Lei Chen;J. Yu

文献摘要

被引文献

相似文献

现有的概率推理查询解决方案主要集中于回答单个推理查询,而很少解决在真实的应用中更为流行和实用的一系列频繁查询的高效返回问题。本文主要研究数据库中推理查询序列之间的计算缓存和共享问题。团树传播(CTP)算法首次引入数据库中的概率推理查询。我们使用物化视图来缓存前一个推理查询的中间结果,这些中间结果可以与后面的查询共享,从而减少了时间开销。此外,我们考虑到查询工作量,以确定频繁查询的变量。为了优化CTP的概率推理查询,我们将这些频繁的查询变量缓存到物化视图中以最大化重用。由于存在不同的查询计划,我们提出了算法来估计成本,并选择最佳的查询计划。最后,我们提出了在关系数据库中的实验评估,以说明我们的方法在回答频繁的概率推理查询的有效性和优越性。
Existing solutions for probabilistic inference queries mainly focus on answering a single inference query, but seldom address the issues of efficiently returning results for a sequence of frequent queries, which is more popular and practical in many real applications. In this paper, we mainly study the computation caching and sharing among a sequence of inference queries in databases. The clique tree propagation (CTP) algorithm is first introduced in databases for probabilistic inference queries. We use the materialized views to cache the intermediate results of the previous inference queries, which might be shared with the following queries, and consequently reduce the time cost. Moreover, we take the query workload into account to identify the frequently queried variables. To optimize probabilistic inference queries with CTP, we cache these frequent query variables into the materialized views to maximize the reuse. Due to the existence of different query plans, we present heuristics to estimate costs and select the optimal query plan. Finally, we present the experimental evaluation in relational databases to illustrate the validity and superiority of our approaches in answering frequent probabilistic inference queries.