Automatic Scatterplot Design Optimization for Clustering Identification

Automatic Scatterplot Design Optimization for Clustering Identification
复制标题

DOI:
10.1109/tvcg.2022.3189883
复制
发表时间:
2022-07
影响因子:
5.2
通讯作者:
Ghulam Jilani Quadri;Jennifer Adorno Nieves;Brenton M. Wiernik;P. Rosen
Ghulam Jilani Quadri;Jennifer Adorno Nieves;Brenton M. Wiernik;P. Rosen
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ghulam Jilani Quadri;Jennifer Adorno Nieves;Brenton M. Wiernik;P. Rosen

文献摘要

相似文献

散点图是最广泛使用的可视化技术之一。引人注目的散点图可视化通过利用视觉感知来提高执行特定视觉分析任务时的意识,从而提高对数据的理解。散点图中的设计选择,如图形编码或数据方面,可以直接影响底层任务(如聚类)的决策质量。因此,构建既考虑视觉编码的感知又考虑正在执行的任务的框架能够优化可视化以最大化功效。在这篇文章中,我们提出了一个自动工具来优化散点图的设计因素,以揭示最显着的集群结构。我们的方法利用合并树数据结构来识别聚类,并优化用于生成散点图图像的子采样算法、采样率、标记大小和标记不透明度的选择。我们通过用户和案例研究验证了我们的方法,表明它可以有效地从大参数空间提供高质量的散点图设计。
Scatterplots are among the most widely used visualization techniques. Compelling scatterplot visualizations improve understanding of data by leveraging visual perception to boost awareness when performing specific visual analytic tasks. Design choices in scatterplots, such as graphical encodings or data aspects, can directly impact decision-making quality for low-level tasks like clustering. Hence, constructing frameworks that consider both the perceptions of the visual encodings and the task being performed enables optimizing visualizations to maximize efficacy. In this article, we propose an automatic tool to optimize the design factors of scatterplots to reveal the most salient cluster structure. Our approach leverages the merge tree data structure to identify the clusters and optimize the choice of subsampling algorithm, sampling rate, marker size, and marker opacity used to generate a scatterplot image. We validate our approach with user and case studies that show it efficiently provides high-quality scatterplot designs from a large parameter space.