Advanced SQL modeling in RDBMS

Advanced SQL modeling in RDBMS
复制标题

RDBMS 中的高级 SQL 建模

DOI:
--
复制
发表时间:
2005
期刊:
TODS
影响因子:
--
通讯作者:
Sankar Subramanian
Sankar Subramanian
中科院分区:
--
文献类型:
--
作者:
Andrew Witkowski;Srikanth Bellamkonda;Tolga Bozkaya;Nathan Folkert;Abhinav Gupta;J. Haydu;Lei Sheng;Sankar Subramanian

文献摘要

被引文献

相似文献

商业关系数据库系统缺乏对复杂业务建模的支持。 ANSI SQL不能将关系视为多维阵列,并定义了多个相互关联的公式,即业务建模所需的操作。关系OLAP(ROLAP)应用程序必须使用JOINS,SQL窗口功能,复杂的案例表达式以及模拟枢轴操作的操作员执行此类任务。 SQL中的指定位置进行计算是Select子句,这是极限限制的,迫使用户以嵌套视图,子查询和复杂的连接来生成查询。此外,SQL查询优化器专注于确定有效的联接订单并选择最佳访问方法,并且在很大程度上无视对多个相互关联的公式的优化。迄今为止,对执行方法的研究集中在有效计算数据立方体和立方体压缩的情况下,而不是用于随机,Internow计算的访问结构。这产生了一个差距,该差距已被电子表格和专门的摩洛普发动机填补,擅长于建模公式的规范,但缺乏关系模型的形式主义,很难在大型用户组中协调,表现出可扩展性问题,并且需要复制。工具和RDBM之间的数据。本文介绍了称为SQL电子表格的SQL扩展名,以提供有关复杂建模关系的数组计算。我们提出了有效处理它们的优化,访问结构和执行模型。特别注意以汇编昂贵的操作(例如聚合)的时间优化。此外,ANSI SQL不能在数据和计算之间提供良好的分离,因此无法支持SQL电子表格模型的参数化。我们为SQL提出了两种参数化方法。一个人使用子查询和标量参数化ANSI SQL视图,该视图允许将数据传递到SQL电子表格。另一种方法介绍了SQL电子表格公式的参数化。这支持构建独立的SQL电子表格库。然后,在模型调用时间期间,这些模型将受到SQL电子表格优化。
Commercial relational database systems lack support for complex business modeling. ANSI SQL cannot treat relations as multidimensional arrays and define multiple, interrelated formulas over them, operations which are needed for business modeling. Relational OLAP (ROLAP) applications have to perform such tasks using joins, SQL Window Functions, complex CASE expressions, and the GROUP BY operator simulating the pivot operation. The designated place in SQL for calculations is the SELECT clause, which is extremely limiting and forces the user to generate queries with nested views, subqueries and complex joins. Furthermore, SQL query optimizers are preoccupied with determining efficient join orders and choosing optimal access methods and largely disregard optimization of multiple, interrelated formulas. Research into execution methods has thus far concentrated on efficient computation of data cubes and cube compression rather than on access structures for random, interrow calculations. This has created a gap that has been filled by spreadsheets and specialized MOLAP engines, which are good at specification of formulas for modeling but lack the formalism of the relational model, are difficult to coordinate across large user groups, exhibit scalability problems, and require replication of data between the tool and RDBMS. This article presents an SQL extension called SQL Spreadsheet, to provide array calculations over relations for complex modeling. We present optimizations, access structures, and execution models for processing them efficiently. Special attention is paid to compile time optimization for expensive operations like aggregation. Furthermore, ANSI SQL does not provide a good separation between data and computation and hence cannot support parameterization for SQL Spreadsheets models. We propose two parameterization methods for SQL. One parameterizes ANSI SQL view using subqueries and scalars, which allows passing data to SQL Spreadsheet. Another method presents parameterization of the SQL Spreadsheet formulas. This supports building stand-alone SQL Spreadsheet libraries. These models are then subject to the SQL Spreadsheet optimizations during model invocation time.