DBMS Architecture – the Layer Model and its Evolution

DBMS Architecture – the Layer Model and its Evolution
复制标题

DBMS 架构 – 层模型及其演变

DOI:
--
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
T. Härder
T. Härder
中科院分区:
--
文献类型:
--
作者:
T. Härder

文献摘要

被引文献

相似文献

第一层的机器因此,在最坏的情况下,在访问存储在磁盘上的数据以评估DB请求之前,必须跨越六个正式接口。这种性能关键的观察促使我们重新考虑系统层的数量。一方面,越来越多的层降低了各个层的复杂性,这反过来又促进了系统的发展。另一方面,越来越多的接口被跨越的DB请求执行增加了运行时开销,并通常降低了DBMS优化的潜力。由于采用了理想的层封装,每个服务调用都意味着参数检查、额外的数据传输和更困难的非本地错误处理。例如,请求的数据必须以其特定于层的格式从层到层复制到请求者。以同样的方式,修改的数据必须传播-再次以其特定于层的格式-最终到达磁盘。因此,层数似乎对整个系统性能有重大影响5。最重要的是,跨系统层的数据复制和更新传播应该最小化。然而,我们提出的架构DBMS模型已经是层的复杂性/系统的发展潜力和请求优化/运行时开销之间的妥协,只要足够的数据映射。我们在2.2节中的讨论已经揭示了减少模型分层结构带来的性能损失的方法。基于这些观察,我们提出了运行时优化我们的静态五层模型,导致两个或三个有效的层在动态的情况下。如图2所示,L5被访问模块取代,访问模块的操作直接引用L4接口。另一方面,使用大的DB缓冲区有效地减少了磁盘5。对于数据库管理系统来说,这一点尤其正确:“性能不是一切,但没有性能,一切都一文不值。·访问,使得几乎所有的逻辑页引用都可以由L2在存储器中定位。有一些预编译的方法,其中访问模块直接将SQL请求映射到L3接口,以进一步节省L4接口的交叉。因此,查询准备被尽可能地推进。甚至特殊查询也可以用这种方式准备,因为事实证明,生成的访问代码比查询解释更有效。通过这种方式,额外的准备成本(在这种情况下延长了查询响应时间)可以快速摊销,特别是当必须访问大量记录时[Chamberlin et al. 1981 a]。另一方面,一些系统传递查询准备(在程序编译时)并使用-以延长查询响应时间为代价-解释器,可以理解为在运行时替代L5。解释器是一个通用程序,在这种情况下,它接受任何SQL语句作为输入,并立即产生查询结果,从而引用L4接口。混合方法在编译时使用一个准备阶段,但不会走极端-访问模块。它们准备中间查询工件,如查询图或执行计划,并使用特定的解释器,通过调用L4操作(以及低层操作)在运行时“执行”这些工件。2.4装订和信息渠道编译和早期装订是个好主意。当然,在处理编译/优化问题时,必须注意重要的方面。缺少运行时参数值引入图2:运行时L4的DBMS模型
machine of layer i. Hence, in the worst case, six formal interfaces have to be crossed before data stored on disk is reached to evaluate a DB request. This performance-critical observation prompts us to reconsider the number of system layers. On the one hand, a growing number of layers reduces the complexity of the individual layers which, in turn, facilitates system evolution. On the other hand, a growing number of interfaces to be crossed for the DB request execution increases the runtime overhead and generally reduces the DBMS optimization potential. Due to the ideal layer encapsulation, each service invocation implies parameter checking, additional data transport, and more difficult handling of non-local errors. For example, the requested data has to be copied – in its layer-specific format – from layer to layer up to the requestor. In the same way, data modified has to be propagated – again in its layer-specific format – eventually down to the disk. Hence, the number of layers seems to have a major influence on the overall system performance5. Above all, data copying and update propagation should be minimized across the system layers. However, our proposed architectural DBMS model is already a compromise between layer complexity/system evolution potential and request optimization/run-time overhead, as far as adequate data mapping is concerned. Our discussion in section 2.2 already revealed the way to reduce the performance penalty introduced by the layered structure of our model. Based on these observations, we propose run-time optimizations to our static five-layer model leading to two or three effective layers in the dynamic case. As illustrated in Fig. 2, L5 is replaced by the access module whose operations directly refer to the L4 interface. At the other side, the use of a large DB buffer effectively reduces disk 5. For DBMSs, it is especially true: »Performance is not everything, but without performance everything is worth nothing.« accesses such that almost all logical page references can be located by L2 in memory. There are precompilation approaches conceivable where the access module directly maps the SQL requests to the L3 interface to further save the crossing of the L4 interface. Hence, query preparation is pushed as far as possible. Even ad-hoc queries may be prepared in this way, because it turned out that generated access code is more effective than query interpretation. In this way, the extra cost of preparation (which extends the query response time in this case) is quickly amortized especially when large sets of records have to be accessed [Chamberlin et al. 1981a]. On the other hand, some systems pass on query preparation (at program compile time) and use – at the cost of extending the query response time – interpreters which can be understood as replacements of L5 at run time. An interpreter is a general program which, in this case, accepts any SQL statement as input and immediately produces its query result thereby referring to the L4 interface. Mixed approaches use a preparation phase at compile time, but do not go to the extreme – the access module. They prepare intermediate query artifacts such as query graphs or execution plans in varying details, and use specific interpreters that »execute« these artifacts at run time by invoking L4 operations (and, in turn, lower-layer operations). 2.4 Binding and Information Channels Compilation and early binding are good ideas. Of course, important aspects have to be observed when dealing with the compilation/optimization problem. Lack of run-time parameter values introduces Fig. 2: DBMS model at run time L4