Unsupervised nearest neighbor regression for dimensionality reduction

Unsupervised nearest neighbor regression for dimensionality reduction
复制标题

DOI:
10.1007/s00500-014-1354-1
复制
发表时间:
2014-07
期刊:
影响因子:
4.1
通讯作者:
Oliver Kramer
Oliver Kramer
中科院分区:
计算机科学3区
文献类型:
--
作者:
Oliver Kramer

文献摘要

被引文献

相似文献

从天文学到生物信息学,各种学科收集了大量的高维模式。在本文中,我们提出了一种基于将最近邻回归拟合到用于学习低维流形的无监督回归框架的非线性降维方法。对于每个高维模式,都会生成一个低维潜在点。引发的优化问题的维数随着模式数量的增加而增加。为了应对较大的解空间,提出了迭代解构建方案。在本文中,我们介绍了两种嵌入高维数据的策略。首先,潜在排序方法允许嵌入到与高维模式的排序相对应的一维潜在空间中。其次,高斯嵌入基于使用数据空间上的距离作为方差的高斯分布采样来随机生成候选位置。核函数通过将模式映射到特征空间来增加该方法的灵活性。我们在一组测试函数上对算法进行实验分析和比较。
Large numbers of high-dimensional patterns are collected in a variety of disciplines, from astronomy to bioinformatics. In this article, we present an approach to non-linear dimensionality reduction based on fitting nearest neighbor regression to the unsupervised regression framework for learning of low-dimensional manifolds. For each high-dimensional pattern, a low-dimensional latent point is generated. The dimensionality of the induced optimization problem grows with the number of patterns. To cope with the large solution space, an iterative solution construction scheme is proposed. In this paper, we introduce two strategies to embed high-dimensional data. First, the latent sorting approach allows embeddings in a one-dimensional latent space corresponding to a sorting of the high-dimensional patterns. Second, Gaussian embeddings randomly generate candidate positions based on sampling from the Gaussian distribution employing distances on data space as variances. Kernel functions increase the flexibility of the approach by mapping the patterns to feature spaces. We analyze and compare the algorithms experimentally on a set of test functions.