Supervised Dimensionality Reduction and Visualization using Centroid-encoder

Supervised Dimensionality Reduction and Visualization using Centroid-encoder
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
ArXiv
影响因子:
--
通讯作者:
T. Ghosh;M. Kirby
T. Ghosh;M. Kirby
中科院分区:
其他
文献类型:
--
作者:
T. Ghosh;M. Kirby

文献摘要

相似文献

可视化高维数据是数据科学和机器学习中的一项重要任务。质心编码器(CE)方法类似于自动编码器,但合并了标签信息,以使类的对象在简化的可视化空间中紧密相连。CE利用非线性和标签来编码低维的高方差,同时捕获数据的全局结构。我们使用各种数据集对该方法进行了详细的分析,并将其与其他监督降维技术进行了比较,包括NCA、非线性NCA、t分布NCA、t分布MCML、监督UMAP、监督PCA、彩色最大方差展开、监督Isomap、参数嵌入、监督邻居检索可视化器和多重关系嵌入。我们的经验表明,质心编码器优于大多数这些技术。我们还表明,当数据方差分布在多个模态时,质心编码器从低维空间的数据中提取了大量信息。这个关键特性确立了它作为数据可视化工具的价值。
Visualizing high-dimensional data is an essential task in Data Science and Machine Learning. The Centroid-Encoder (CE) method is similar to the autoencoder but incorporates label information to keep objects of a class close together in the reduced visualization space. CE exploits nonlinearity and labels to encode high variance in low dimensions while capturing the global structure of the data. We present a detailed analysis of the method using a wide variety of data sets and compare it with other supervised dimension reduction techniques, including NCA, nonlinear NCA, t-distributed NCA, t-distributed MCML, supervised UMAP, supervised PCA, Colored Maximum Variance Unfolding, supervised Isomap, Parametric Embedding, supervised Neighbor Retrieval Visualizer, and Multiple Relational Embedding. We empirically show that centroid-encoder outperforms most of these techniques. We also show that when the data variance is spread across multiple modalities, centroid-encoder extracts a significant amount of information from the data in low dimensional space. This key feature establishes its value to use it as a tool for data visualization.