Trinity: A Fast Compressed Multi-attribute Data Store

Trinity: A Fast Compressed Multi-attribute Data Store
复制标题

DOI:
10.1145/3627703.3650072
复制
发表时间:
2024-04
期刊:
Proceedings of the Nineteenth European Conference on Computer Systems
影响因子:
--
通讯作者:
Ziming Mao;K. Srinivasan;Anurag Khandelwal
Ziming Mao;K. Srinivasan;Anurag Khandelwal
中科院分区:
其他
文献类型:
--
作者:
Ziming Mao;K. Srinivasan;Anurag Khandelwal

文献摘要

相似文献

随着属性丰富的机器生成数据的激增,新兴的实时监控、诊断和可视化工具同时跨多个属性摄取和分析此类数据。由于数据量巨大,应用程序需要高效的存储和高性能的数据表示来有效地分析它们。我们提出Trinity,一个系统,同时促进查询和存储效率在大量的多属性记录。Trinity通过一种新的动态、简洁、多维的数据结构MdTrie来实现这一点。MdTrie采用了新的莫顿码泛化,多属性查询算法和自索引Trie结构的组合,以实现上述目标。我们对Trinity的真实用例评估表明,与最先进的系统相比,它支持(1)7.2-59.6倍的多属性搜索速度,(2)与OLAP列存储相当的存储占用空间,比NoSQL存储和OLTP数据库低4.8-15.1倍,点查询吞吐量与NoSQL存储相当,比OLTP数据库和OLAP列式存储高1.7-52.5倍。
With the proliferation of attribute-rich machine-generated data, emerging real-time monitoring, diagnosis, and visualization tools ingest and analyze such data across multiple attributes simultaneously. Due to the sheer volume of the data, applications need storage-efficient and performant data representations to analyze them efficiently. We present TRINITY, a system that simultaneously facilitates query and storage efficiency across large volumes of multi-attribute records. Trinity accomplishes this through a new dynamic, succinct, multi-dimensional data structure, MdTrie. MdTrie employs a combination of novel Morton code generalization, a multi-attribute query algorithm, and a self-indexed trie structure to achieve the above goals. Our evaluation of TRINITY for real-world use-cases shows that compared to state-of-the-art systems, it supports (1) 7.2-59.6× faster multi-attribute searches, (2) storage footprint comparable to OLAP columnar stores and 4.8-15.1× lower than NoSQL stores and OLTP databases, and (3) point query throughput comparable to NoSQL stores and 1.7-52.5× higher than OLTP databases and OLAP columnar stores.