CAREER: Scrapple: Fast Analytical Query Evaluation via Advanced Query Recycling Techniques
CAREER: Scrapple: Fast Analytical Query Evaluation via Advanced Query Recycling Techniques
批准号:
1055107
负责人:
Todd Green
金额:
$55.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-01-01 至 2014-10-31
中文摘要
以决策支持应用程序为特征的复杂分析查询的计算成本可能非常高,并且此类应用程序的价值与向用户返回答案的速度直接相关。通常,一旦查询得到回答,数据库系统就会简单地丢弃结果。然而,这样做会错过一个巨大的优化机会:如果我们只知道如何回收查询结果来帮助回答后续的相关查询,那么丢弃的查询结果中存在着巨大的潜力。该项目的目标是开发Scrapple,这是一个原则性的数据库管理系统,它积极地重用旧的查询结果来加快对新查询的回答,从而为一大类决策支持应用程序带来潜在的显著性能提升。Scrapple的基本策略是将缓存的查询结果(及其中间子结果)作为物化视图查看,然后使用高级技术优化查询,使用物化视图来回答后续查询。为了执行这一策略,该项目开发了:(1)一种新颖而全面的差异重构策略理论;(2)一套连接增量视图维护和使用物化视图优化查询的统一原则;(3)一种新颖而全面的聚合查询数据来源理论;以及(4)通过基于成本的搜索策略回收缓存结果的实用实现技术。通过使用全自动化技术,Scrapple将极大地降低典型数据仓库的总拥有成本。此外,我们方法的核心技术在数据集成、数据交换、视图维护和数据起源等领域有广泛的应用。这项研究还将被用来为新的课程模块开发讲座和项目材料。这些教育材料,连同SCRAPPLE源代码和出版物,将在项目网站http://www.cs.ucdavis.edu/~green/scrapple.上免费提供
英文摘要
The complex analytical queries characterizing decision support applications can be very expensive to compute, and the value of such applications is directly correlated to the speed at which answers can be returned to the user. Typically, once queries have been answered, database systems simply discard the results. However, a huge optimization opportunity is missed by doing this: there is tremendous latent energy in the discarded query results, if we only knew how to recycle them to help answer subsequent related queries. The goal of the project is to develop Scrapple, a principled database management system that aggressively reuses old query results to speed up the answering of new queries, resulting in potentially dramatic performance gains for a large class of decision support applications.Scrapple's basic strategy is to view cached query results (and their intermediate subresults) as materialized views, and then employ advanced techniques for optimizing queries using materialized views to answer subsequent queries. To execute this strategy, the project develops: (1) a novel and comprehensive theory of differential reformulation strategies; (2) a set of unifying principles connecting incremental view maintenance and optimization of queries using materialized views; (3) a novel and comprehensive theory of data provenance for aggregate queries; and (4) practical implementation techniques for recycling cached results via cost-based search strategies. By using fully automated techniques, Scrapple will dramatically reduce the total cost of ownership of a typical data warehouse. Moreover, the techniques at the heart of our approach have wide application in areas such as data integration, data exchange, view maintenance, and data provenance. The research will also be used to develop lecture and project materials for new course modules. These educational materials, along the Scrapple source code and publications, will be made freely available at the project Web site, http://www.cs.ucdavis.edu/~green/scrapple.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文