Nanocubes for Real-Time Exploration of Spatiotemporal Datasets

Nanocubes for Real-Time Exploration of Spatiotemporal Datasets
复制标题

DOI:
10.1109/tvcg.2013.179
复制
发表时间:
2013-12-01
影响因子:
5.2
通讯作者:
Scheidegger, Carlos
Scheidegger, Carlos
中科院分区:
计算机科学1区
文献类型:
--
作者:
Lins, Lauro;Klosowski, James T.;Scheidegger, Carlos

文献摘要

被引文献

相似文献

考虑对具有数十亿个条目的大型多维时空数据集的实时探索,每个条目由位置、时间和其他属性定义。某些属性在空间上或时间上是相关的?数据中有没有趋势或离群值?回答这些问题需要对域的任意区域和数据的属性进行聚合。许多关系数据库实现众所周知的数据立方体聚合操作,在某种意义上,该操作预先计算数据库上的每个可能的聚合查询。数据立方体有时被认为占用了令人望而却步的大量空间,因此需要磁盘存储。相反,我们展示了如何构建适合现代笔记本电脑主内存的数据立方体,甚至可以容纳数十亿个条目;我们将这种数据结构称为纳米立方体。我们给出了计算和查询纳米立方体的算法,并展示了如何使用它来生成众所周知的可视编码,如热图、直方图和平行坐标图。与通过扫描整个数据集创建的精确可视化相比,由于空间和时间上的分层结构,纳米立方体图在各种尺度上具有有限的屏幕误差。我们在各种真实数据集上演示了我们的技术的有效性,并提供了内存、计时和网络带宽测量。我们发现,在我们的示例中,查询的时间主要受网络和用户交互延迟的影响。
Consider real-time exploration of large multidimensional spatiotemporal datasets with billions of entries, each defined by a location, a time, and other attributes. Are certain attributes correlated spatially or temporally? Are there trends or outliers in the data? Answering these questions requires aggregation over arbitrary regions of the domain and attributes of the data. Many relational databases implement the well-known data cube aggregation operation, which in a sense precomputes every possible aggregate query over the database. Data cubes are sometimes assumed to take a prohibitively large amount of space, and to consequently require disk storage. In contrast, we show how to construct a data cube that fits in a modern laptop's main memory, even for billions of entries; we call this data structure a nanocube. We present algorithms to compute and query a nanocube, and show how it can be used to generate well-known visual encodings such as heatmaps, histograms, and parallel coordinate plots. When compared to exact visualizations created by scanning an entire dataset, nanocube plots have bounded screen error across a variety of scales, thanks to a hierarchical structure in space and time. We demonstrate the effectiveness of our technique on a variety of real-world datasets, and present memory, timing, and network bandwidth measurements. We find that the timings for the queries in our examples are dominated by network and user-interaction latencies.