Enabling query processing across heterogeneous data models: A survey

Enabling query processing across heterogeneous data models: A survey
复制标题

DOI:
10.1109/bigdata.2017.8258302
复制
发表时间:
2017-12
期刊:
2017 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Ran Tan;Rada Y. Chirkova;V. Gadepally;T. Mattson
Ran Tan;Rada Y. Chirkova;V. Gadepally;T. Mattson
中科院分区:
其他
文献类型:
--
作者:
Ran Tan;Rada Y. Chirkova;V. Gadepally;T. Mattson

文献摘要

被引文献

相似文献

现代应用程序通常需要管理和分析跨越多个数据模型的各种数据集[1],[2],[3],[4],[5]。在这种情况下,通过提取-转换-加载(ETL)过程存储数据的成本可能很高。将不同的数据转换为单个数据模型可能会降低性能。此外,管理不同的数据集和维护管道可以证明是劳动密集型的。因此,一个新兴的趋势是将重点转移到联合专门的数据存储和支持跨异构数据模型的查询处理[6]。这种转变可以带来许多好处:首先,系统可以原生地利用多个数据模型,这可以转化为最大化底层接口的语义表达能力,并利用组件数据存储的内部处理能力。第二,联邦体系结构支持特定于查询的数据集成和即时转换和迁移,这有可能显著降低操作复杂性和开销。侧重于开发这一研究领域系统的项目来自不同的背景,并解决不同的问题,这可能使人们难以对这一领域的工作形成一致的看法。在这项调查中,我们介绍了一个分类描述的最先进的状态,并提出了一个系统的评估框架,有利于了解查询处理的相关系统的特点。我们使用该框架来评估四个代表性的实现:BigDAWG [7],[8],CloudMdsQL [9],[10],Myria [11],[12]和Apache Drill [13]。
Modern applications often need to manage and analyze widely diverse datasets that span multiple data models [1], [2], [3], [4], [5]. Warehousing the data through Extract-Transform-Load (ETL) processes can be expensive in such scenarios. Transforming disparate data into a single data model may degrade performance. Further, curating diverse datasets and maintaining the pipeline can prove to be labor intensive. As a result, an emerging trend is to shift the focus to federating specialized data stores and enabling query processing across heterogeneous data models [6]. This shift can bring many advantages: First, systems can natively leverage multiple data models, which can translate to maximizing the semantic expressiveness of underlying interfaces and leveraging the internal processing capabilities of component data stores. Second, federated architectures support query-specific data integration with just-in-time transformation and migration, which has the potential to significantly reduce the operational complexity and overhead. Projects that focus on developing systems in this research area stem from various backgrounds and address diverse concerns, which could make it difficult to form a consistent view of the work in this area. In this survey, we introduce a taxonomy for describing the state of the art and propose a systematic evaluation framework conducive to understanding of query-processing characteristics in the relevant systems. We use the framework to assess four representative implementations: BigDAWG [7], [8], CloudMdsQL [9], [10], Myria [11], [12], and Apache Drill [13].