Making many-to-many parallel coordinate plots scalable by asymmetric biclustering

Making many-to-many parallel coordinate plots scalable by asymmetric biclustering
复制标题

DOI:
10.1109/pacificvis.2017.8031609
复制
发表时间:
2017-09
期刊:
2017 IEEE Pacific Visualization Symposium (PacificVis)
影响因子:
--
通讯作者:
Hsiang-Yun Wu;Yusuke Niibe;Kazuho Watanabe;Shigeo Takahashi;M. Uemura;I. Fujishiro
Hsiang-Yun Wu;Yusuke Niibe;Kazuho Watanabe;Shigeo Takahashi;M. Uemura;I. Fujishiro
中科院分区:
其他
文献类型:
--
作者:
Hsiang-Yun Wu;Yusuke Niibe;Kazuho Watanabe;Shigeo Takahashi;M. Uemura;I. Fujishiro

文献摘要

相似文献

通过最近先进的测量技术获得的数据集往往具有大量的维度。这导致分析这些数据集的计算成本急剧增加,从而使科学假设的制定和验证变得非常困难。因此,需要一种有效的方法来识别目标数据集的特征子空间,即数据样本的维度变量或子集的子空间,以描述原始数据集中隐藏的本质。本文提出了一种支持半自动数据分析的可视化数据挖掘框架,该框架基于非对称双聚类来探索高度相关的特征子空间。为此,本文扩展了并行坐标图的一种变体——多对多并行坐标图,以在视觉上帮助适当地选择特征子空间,并避免固有的视觉混乱。在该框架中,对数据集的维度变量和数据样本同时非对称地进行双聚类。将一组可变轴投影到单个复合轴上,而将两个连续可变轴之间的数据样本使用多边形条进行捆绑。这使得可视化方法具有可伸缩性,并使其能够在框架中发挥关键作用。该框架的有效性已被经验证明,它对多对多平行坐标图非常有用。
Datasets obtained through recently advanced measurement techniques tend to possess a large number of dimensions. This leads to explosively increasing computation costs for analyzing such datasets, thus making formulation and verification of scientific hypotheses very difficult. Therefore, an efficient approach to identifying feature subspaces of target datasets, that is, the subspaces of dimension variables or subsets of the data samples, is required to describe the essence hidden in the original dataset. This paper proposes a visual data mining framework for supporting semiautomatic data analysis that builds upon asymmetric biclustering to explore highly correlated feature subspaces. For this purpose, a variant of parallel coordinate plots, many-to-many parallel coordinate plots, is extended to visually assist appropriate selections of feature subspaces as well as to avoid intrinsic visual clutter. In this framework, biclustering is applied to dimension variables and data samples of the dataset simultaneously and asymmetrically. A set of variable axes are projected to a single composite axis while data samples between two consecutive variable axes are bundled using polygonal strips. This makes the visualization method scalable and enables it to play a key role in the framework. The effectiveness of the proposed framework has been empirically proven, and it is remarkably useful for many-to-many parallel coordinate plots.