Derby/S: a DBMS for sample-based query answering

Derby/S: a DBMS for sample-based query answering
复制标题

Derby/S:用于基于样本的查询应答的 DBMS

DOI:
--
复制
发表时间:
2006
期刊:
SIGMOD Conference
影响因子:
--
通讯作者:
Wolfgang Lehner
Wolfgang Lehner
中科院分区:
--
文献类型:
--
作者:
Anja Klein;Rainer Gemulla;Philipp J. Rösch;Wolfgang Lehner

文献摘要

被引文献

相似文献

虽然近似查询处理是满足数据分析应用需求的一种重要方式,但目前的数据库系统并没有为这些技术提供集成和全面的支持。为了改善这种情况,我们提出了一个SQL扩展(称为SQL/S),用于使用随机样本进行近似查询应答,并在开源数据库系统Derby(称为Derby/S)的引擎中提供了一个原型实现。我们的方法通过以声明的方式定义样本,显着减少了所需的专家知识;具体采样方案及其参数化的选择由系统自行决定。SQL/S引入了新的DDL命令,可以根据一组给定的优化标准轻松定义和管理随机样本。如果底层数据集发生变化,Derby/S将自动负责样例维护。最后,在查询处理过程中透明地使用样本,并提供错误边界。我们的扩展不影响传统查询,并提供了将采样作为头等公民集成到DBMS中的方法。
Although approximate query processing is a prominent way to cope with the requirements of data analysis applications, current database systems do not provide integrated and comprehensive support for these techniques. To improve this situation, we propose an SQL extension---called SQL/S---for approximate query answering using random samples, and present a prototypical implementation within the engine of the open-source database system Derby---called Derby/S. Our approach significantly reduces the required expert knowledge by enabling the definition of samples in a declarative way; the choice of the specific sampling scheme and its parametrization is left to the system. SQL/S introduces new DDL commands to easily define and administrate random samples subject to a given set of optimization criteria. Derby/S automatically takes care of sample maintenance if the underlying dataset changes. Finally, samples are transparently used during query processing, and error bounds are provided. Our extensions do not affect traditional queries and provide the means to integrate sampling as a first-class citizen into a DBMS.