STASH : Fast Hierarchical Aggregation Queries for Effective Visual Spatiotemporal Explorations
STASH : Fast Hierarchical Aggregation Queries for Effective Visual Spatiotemporal Explorations
复制标题
DOI:
10.1109/cluster.2019.8891029
复制
发表时间:
2019-09
期刊:
影响因子:
--
通讯作者:
Saptashwa Mitra;Paahuni Khandelwal;S. Pallickara;S. Pallickara
中科院分区:
文献类型:
--
作者:
Saptashwa Mitra;Paahuni Khandelwal;S. Pallickara;S. Pallickara
The proliferation of sensors and observational instruments enable scientists to explore natural, spatiotemporal phenomena via explorative analysis and advanced modeling. Geospatial visualization, in particular, is an intuitive tool to identify patterns, enhance understanding of the data, and plan for subsequent analysis. However, seamless interactions between end-user devices and the sheer volume of data have been a challenge due to the limited bandwidth and data access latencies.In this paper, we introduce Stash, a distributed, in-memory cache for hierarchical aggregation and query evaluations. Stash is a middleware which can be loaded on top of a distributed file system. Users perform queries from a lightweight visualization interface at the front-end and the evaluations occur over the back-end storage system housing the raw data over which summarization and subsequent visualizations are to be performed. Stash facilitates fast exploratory analytics by caching relevant past query results based on their frequency and freshness to assist similar, future queries and avoid expensive disk I/O and network usage, thus reducing their latency. Additionally, Stash handles any hotspot that might result from a spike in user requests due to the spatial and temporal locality of their access patterns.Our empirical benchmarks show that a Stash-enabled system reduces query latency of a basic system by over 5-folds and brings it down to interactive speed even for large country-sized spatiotemporal queries. We have contrasted Stash with existing cache-enabled analytics engines, such as ElasticSearch, and found that our STASH-enabled system reduced the aggregation query latency up to ~70%. STASH also alleviated skewed workloads through its dynamic replication scheme and improved throughput by ~40% in hotspot scenarios.