Polygonal Coordinate System: Visualizing high-dimensional data using geometric DR, and a deterministic version of t-SNE

Polygonal Coordinate System: Visualizing high-dimensional data using geometric DR, and a deterministic version of t-SNE
复制标题

DOI:
10.1016/j.eswa.2021.114741
复制
发表时间:
2021-03-31
影响因子:
8.5
通讯作者:
Sales, Claudomiro
Sales, Claudomiro
中科院分区:
计算机科学1区
文献类型:
--
作者:
Flexa, Caio;Gomes, Walisson;Sales, Claudomiro

文献摘要

被引文献

相似文献

抽象性约简(DR)对于理解高维数据是非常有用的。它吸引了工业界和学术界的广泛关注,并在机器学习,数据挖掘和模式识别等领域得到应用。这项工作提出了一种几何方法DR称为多边形坐标系(PCS),能够表示多维数据在两个或三个维度,同时保留其固有的整体结构,利用多边形接口桥接高,低维空间。PCS可以通过采用线性时间复杂度的增量几何DR来处理大数据。还提供了一种新版本的t-分布随机邻居嵌入(t-SNE),一种最先进的DR算法。它采用基于PCS的确定性策略,并被命名为t-分布式确定性邻居嵌入(t-DNE)。几个合成和真实的数据集被用作我们的基准测试中众所周知的现实世界的问题原型,提供了一种方法来评估PCS和t-DNE对四个基于嵌入的DR算法:两个线性变换的(主成分分析和非负矩阵分解)和两个非线性的(t-SNE和Sammon映射)。这些算法的执行时间的统计比较,弗里德曼的显着性测试,突出了PCS在数据嵌入的效率。PCS在本工作中探索的几个方面往往超过其同行,包括渐近时间和空间复杂性,全局数据固有结构的保留,超参数的数量以及对未观测数据的适用性。
Dimensionality Reduction (DR) is useful to understand high-dimensional data. It attracts wide attention from industry and academia and is employed in areas such as machine learning, data mining, and pattern recognition. This work presents a geometric approach to DR termed Polygonal Coordinate System (PCS), capable of representing multidimensional data in two or three dimensions while preserving their inherent overall structure by taking advantage of a polygonal interface bridging high-and low-dimensional spaces. PCS can handle Big Data by adopting an incremental, geometric DR with linear-time complexity. A new version of t-Distributed Stochastic Neighbor Embedding (t-SNE), a state-of-the-art algorithm for DR, is also provided. It employs a PCS-based deterministic strategy and is named t-Distributed Deterministic Neighbor Embedding (t-DNE). Several synthetic and real data sets were used as well-known real-world problem archetypes in our benchmark, providing a means to evaluate PCS and t-DNE against four embedding-based DR algorithms: two linear-transformation ones (Principal Component Analysis and Non-negative Matrix Factorization) and two nonlinear ones (t-SNE and Sammon's Mapping). Statistical comparisons of the execution times of these algorithms, by the Friedman's significance test, highlight the efficiency of PCS in data embedding. PCS tends to surpass its counterparts in several aspects explored in this work, including asymptotic time and space complexity, preservation of global data-inherent structures, number of hyperparameters, and applicability to unobserved data.