Query Planning for Continuous Aggregation Queries over a Network of Data Aggregators

Query Planning for Continuous Aggregation Queries over a Network of Data Aggregators
复制标题

DOI:
10.1109/tkde.2011.12
复制
发表时间:
2012-06
影响因子:
8.9
通讯作者:
Rajeev Gupta;K. Ramamritham
Rajeev Gupta;K. Ramamritham
中科院分区:
计算机科学2区
文献类型:
--
作者:
Rajeev Gupta;K. Ramamritham

文献摘要

被引文献

相似文献

连续查询用于监视随时间变化的数据的变化,并提供对在线决策有用的结果。通常,用户希望获得分布式数据项上的某些聚合函数的值,例如,了解客户的投资组合的价值;或一组传感器感测到的温度的平均值。在这些查询中,客户端指定一致性要求作为查询的一部分。我们提出了一种低成本、可扩展的技术,使用动态数据项聚合器网络来回答连续聚合查询。在这样的数据聚合器网络中,每个数据聚合器以特定的一致性提供一组数据项。正如动态网页的各个片段由内容分发网络的一个或多个节点提供服务一样,我们的技术涉及将客户端查询分解为子查询,并在明智选择的数据聚合器上执行子查询,并具有各自的子查询不一致边界。我们提供了一种技术,用于获取具有不一致性边界的最佳子查询集,该技术以最少数量的从聚合器发送到客户端的刷新消息来满足客户端查询的一致性要求。为了估计刷新消息的数量,我们构建了一个查询成本模型,该模型可用于估计满足客户端指定的不一致性界限所需的消息数量。使用真实世界跟踪的性能结果表明,我们基于成本的查询规划导致执行查询时使用的消息数量不到现有方案所需消息数量的三分之一。
Continuous queries are used to monitor changes to time varying data and to provide results useful for online decision making. Typically a user desires to obtain the value of some aggregation function over distributed data items, for example, to know value of portfolio for a client; or the AVG of temperatures sensed by a set of sensors. In these queries a client specifies a coherency requirement as part of the query. We present a low-cost, scalable technique to answer continuous aggregation queries using a network of aggregators of dynamic data items. In such a network of data aggregators, each data aggregator serves a set of data items at specific coherencies. Just as various fragments of a dynamic webpage are served by one or more nodes of a content distribution network, our technique involves decomposing a client query into subqueries and executing subqueries on judiciously chosen data aggregators with their individual subquery incoherency bounds. We provide a technique for getting the optimal set of subqueries with their incoherency bounds which satisfies client query's coherency requirement with least number of refresh messages sent from aggregators to the client. For estimating the number of refresh messages, we build a query cost model which can be used to estimate the number of messages required to satisfy the client specified incoherency bound. Performance results using real-world traces show that our cost-based query planning leads to queries being executed using less than one third the number of messages required by existing schemes.