Supporting Analysis of Dimensionality Reduction Results with Contrastive Learning

Supporting Analysis of Dimensionality Reduction Results with Contrastive Learning
复制标题

DOI:
10.1109/tvcg.2019.2934251
复制
发表时间:
2020-01-01
影响因子:
5.2
通讯作者:
Ma, Kwan-Liu
Ma, Kwan-Liu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Fujiwara, Takanori;Kwon, Oh-Hyun;Ma, Kwan-Liu

文献摘要

被引文献

相似文献

简化(DR)经常用于分析和可视化高维数据,因为它提供了良好的数据初视图。然而,为了解释DR结果以从数据中获得有用的见解,需要进行额外的分析工作,例如识别聚类并了解其特征。虽然有许多自动方法(例如,基于密度的聚类方法)来识别聚类,但是仍然缺乏用于理解聚类特征的有效方法。聚类的主要特征在于其特征值的分布。当特征数量很大时,检查原始特征值不是一项简单的任务。为了应对这一挑战,我们提出了一种可视化的分析方法,有效地突出了灾难恢复结果中集群的基本特征。为了提取基本特征,我们引入了对比主成分分析(cPCA)的增强用法。我们的方法,称为ccPCA(对比聚类PCA),可以计算每个特征对一个聚类和其他聚类之间对比度的相对贡献。使用ccPCA,我们创建了一个交互式系统,包括聚类特征贡献的可扩展可视化。我们证明了我们的方法和系统的有效性,使用几个公开的数据集的案例研究。
Dimensionality reduction (DR) is frequently used for analyzing and visualizing high-dimensional data as it provides a good first glance of the data. However, to interpret the DR result for gaining useful insights from the data, it would take additional analysis effort such as identifying clusters and understanding their characteristics. While there are many automatic methods (e.g., density-based clustering methods) to identify clusters, effective methods for understanding a clusters characteristics are still lacking. A cluster can be mostly characterized by its distribution of feature values. Reviewing the original feature values is not a straightforward task when the number of features is large. To address this challenge, we present a visual analytics method that effectively highlights the essential features of a cluster in a DR result. To extract the essential features, we introduce an enhanced usage of contrastive principal component analysis (cPCA). Our method, called ccPCA (contrasting clusters in PCA), can calculate each features relative contribution to the contrast between one cluster and other clusters. With ccPCA, we have created an interactive system including a scalable visualization of clusters feature contributions. We demonstrate the effectiveness of our method and system with case studies using several publicly available datasets.