STASH : Fast Hierarchical Aggregation Queries for Effective Visual Spatiotemporal Explorations

STASH : Fast Hierarchical Aggregation Queries for Effective Visual Spatiotemporal Explorations
复制标题

DOI:
10.1109/cluster.2019.8891029
复制
发表时间:
2019-09
期刊:
2019 IEEE International Conference on Cluster Computing (CLUSTER)
影响因子:
--
通讯作者:
Saptashwa Mitra;Paahuni Khandelwal;S. Pallickara;S. Pallickara
Saptashwa Mitra;Paahuni Khandelwal;S. Pallickara;S. Pallickara
中科院分区:
其他
文献类型:
--
作者:
Saptashwa Mitra;Paahuni Khandelwal;S. Pallickara;S. Pallickara

文献摘要

被引文献

相似文献

传感器和观测仪器的激增使科学家能够通过探索性分析和高级建模来探索自然、时空现象。尤其是地理空间可视化,它是一种直观的工具,用于识别模式、增强对数据的理解和为后续分析做计划。然而,由于有限的带宽和数据访问延迟,最终用户设备与大量数据之间的无缝交互一直是一个挑战。在本文中,我们介绍了Stash,一个用于分层聚合和查询计算的分布式内存缓存。Stash是一个中间件,可以加载在分布式文件系统的顶部。用户从前端的轻量级可视化界面执行查询,并在保存原始数据的后端存储系统上执行评估,汇总和后续可视化将在原始数据上执行。Stash通过缓存过去相关查询结果的频率和新鲜度来促进快速的探索性分析,以帮助类似的、未来的查询,避免昂贵的磁盘I/O和网络使用,从而减少延迟。此外,Stash还可以处理由于用户访问模式的空间和时间局域性而导致的用户请求激增所导致的热点。我们的经验基准测试表明,启用隐藏的系统将基本系统的查询延迟减少了5倍以上,并将其降低到交互速度,即使对于大型国家/地区的时空查询也是如此。我们将Stash与现有的支持缓存的分析引擎(如ElasticSearch)进行了对比,发现我们的支持缓存的系统将聚合查询延迟减少了约70%。STASH还通过其动态复制方案减轻了倾斜的工作负载,并在热点场景中将吞吐量提高了约40%。
The proliferation of sensors and observational instruments enable scientists to explore natural, spatiotemporal phenomena via explorative analysis and advanced modeling. Geospatial visualization, in particular, is an intuitive tool to identify patterns, enhance understanding of the data, and plan for subsequent analysis. However, seamless interactions between end-user devices and the sheer volume of data have been a challenge due to the limited bandwidth and data access latencies.In this paper, we introduce Stash, a distributed, in-memory cache for hierarchical aggregation and query evaluations. Stash is a middleware which can be loaded on top of a distributed file system. Users perform queries from a lightweight visualization interface at the front-end and the evaluations occur over the back-end storage system housing the raw data over which summarization and subsequent visualizations are to be performed. Stash facilitates fast exploratory analytics by caching relevant past query results based on their frequency and freshness to assist similar, future queries and avoid expensive disk I/O and network usage, thus reducing their latency. Additionally, Stash handles any hotspot that might result from a spike in user requests due to the spatial and temporal locality of their access patterns.Our empirical benchmarks show that a Stash-enabled system reduces query latency of a basic system by over 5-folds and brings it down to interactive speed even for large country-sized spatiotemporal queries. We have contrasted Stash with existing cache-enabled analytics engines, such as ElasticSearch, and found that our STASH-enabled system reduced the aggregation query latency up to ~70%. STASH also alleviated skewed workloads through its dynamic replication scheme and improved throughput by ~40% in hotspot scenarios.