Small Refinements to the DAM Can Have Big Consequences for Data-Structure Design
Small Refinements to the DAM Can Have Big Consequences for Data-Structure Design
复制标题
对 DAM 的小改进可能会对数据结构设计产生重大影响
DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Yang Zhan
中科院分区:
文献类型:
--
作者:
M. A. Bender;Alex Conway;Martín Farach;William K. Jannen;Yizheng Jiao;Rob Johnson;Eric R. Knorr;Sara McAllister;Nirjhar Mukherjee;P. Pandey;Donald E. Porter;Jun Yuan;Yang Zhan
Storage devices have complex performance profiles, including costs to initiate IOs (e.g., seek times in hard drives), parallelism and bank conflicts (in SSDs), costs to transfer data, and firmware-internal operations. The Disk-Access Machine (DAM) model simplifies reality by assuming that storage devices transfer data in blocks of size B and that all transfers have unit cost. Despite its simplifications, the DAM model is reasonably accurate. In fact, if B is set to the half-bandwidth point, where the latency and bandwidth of the hardware are equal, the DAM approximates the IO cost on any hardware to within a factor of 2. Furthermore, the DAM explains the popularity of B-trees in the 70s and the current popularity of B-trees and log-structured merge trees. But it fails to explain why some B-trees use small nodes, whereas all B-trees use large nodes. In a DAM, all IOs, and hence all nodes, are the same size. In this paper, we show that the affine and PDAM models, which are small refinements of the DAM model, yield a surprisingly large improvement in predictability without sacrificing ease of use. We present benchmarks on a large collection of storage devices showing that the affine and PDAM models give good approximations of the performance characteristics of hard drives and SSDs, respectively. We show that the affine model explains node-size choices in B-trees and B+-trees. Furthermore, the models predict that the B-tree is highly sensitive to variations in the node size whereas B-trees are much less sensitive. These predictions are born out empirically. Finally, we show that in both the affine and PDAM models, it pays to organize data structures to exploit varying IO size. In the affine model, B-trees can be optimized so that all operations are simultaneously optimal, even up to lower order terms. In the PDAM model, B-trees (or B+-trees) can be organized so that both sequential and concurrent workloads are handled efficiently. We conclude that the DAM model is useful as a first cut when designing or analyzing an algorithm or data structure but the affine and PDAM models enable the algorithm designer to optimize parameter choices and fill in design details.
登录
查看更多内容
DOI:
10.1145/2935764.2935767
发表时间:
2016-07
期刊:
Proceedings of the 28th ACM Symposium on Parallelism in Algorithms and Architectures
影响因子:
--
作者:
N. Ben-David;G. Blelloch;Jeremy T. Fineman;Phillip B. Gibbons;Yan Gu;Charles McGuffey;Julian Shun
通讯作者:
N. Ben-David;G. Blelloch;Jeremy T. Fineman;Phillip B. Gibbons;Yan Gu;Charles McGuffey;Julian Shun
DOI:
10.1145/3034786.3056117
发表时间:
2017-05
期刊:
Proceedings of the 36th ACM SIGMOD-SIGACT-SIGAI Symposium on Principles of Database Systems
影响因子:
--
作者:
M. A. Bender;Martín Farach-Colton;Rob Johnson;Simon Mauras;Tyler Mayer;C. Phillips;Helen Xu
通讯作者:
M. A. Bender;Martín Farach-Colton;Rob Johnson;Simon Mauras;Tyler Mayer;C. Phillips;Helen Xu
DOI:
10.1145/3210377.3210381
发表时间:
2018-05
期刊:
Proceedings of the 30th on Symposium on Parallelism in Algorithms and Architectures
影响因子:
--
作者:
G. Blelloch;Phillip B. Gibbons;Yan Gu;Charles McGuffey;Julian Shun
通讯作者:
G. Blelloch;Phillip B. Gibbons;Yan Gu;Charles McGuffey;Julian Shun
DOI:
10.4230/lipics.icalp.2018.39
发表时间:
2018-05
期刊:
--
影响因子:
--
作者:
Alex Conway;Martín Farach-Colton;Philip Shilane
通讯作者:
Alex Conway;Martín Farach-Colton;Philip Shilane