Parametric nonlinear dimensionality reduction using kernel t-SNE

Parametric nonlinear dimensionality reduction using kernel t-SNE
复制标题

DOI:
10.1016/j.neucom.2013.11.045
复制
发表时间:
2015-01-05
期刊:
影响因子:
6
通讯作者:
Hammer, Barbara
Hammer, Barbara
中科院分区:
计算机科学2区
文献类型:
--
作者:
Gisbrecht, Andrej;Schulz, Alexander;Hammer, Barbara

文献摘要

被引文献

相似文献

新的非参数降维技术,如t分布随机邻居嵌入(t-SNE),导致一个强大的和灵活的可视化高维数据。非参数技术的一个缺点是它们缺乏显式的样本外扩展。在这方面的贡献,我们提出了一个有效的扩展t-SNE的参数化框架,内核t-SNE,它保留了基本的t-SNE的灵活性,但使明确的样本外扩展。我们测试的能力内核t-SNE相比,标准的t-SNE基准数据集,特别是解决的泛化能力的映射新的数据。在大数据集的背景下,这个过程使我们能够只为固定大小的子集训练映射,然后在线性时间内映射所有数据。我们证明,这种技术产生令人满意的结果,也为大数据集提供缺失的信息,由于小尺寸的子集占辅助信息,如类标签,它可以集成到内核的t-SNE的基础上的Fisher信息。(C)2014爱思唯尔有限公司版权所有。
Novel non-parametric dimensionality reduction techniques such as t-distributed stochastic neighbor embedding (t-SNE) lead to a powerful and flexible visualization of high-dimensional data. One drawback of non-parametric techniques is their lack of an explicit out-of-sample extension. In this contribution, we propose an efficient extension of t-SNE to a parametric framework, kernel t-SNE, which preserves the flexibility of basic t-SNE, but enables explicit out-of-sample extensions. We test the ability of kernel t-SNE in comparison to standard t-SNE for benchmark data sets, in particular addressing the generalization ability of the mapping for novel data. In the context of large data sets, this procedure enables us to train a mapping for a fixed size subset only, mapping all data afterwards in linear time. We demonstrate that this technique yields satisfactory results also for large data sets provided missing information due to the small size of the subset is accounted for by auxiliary information such as class labels, which can be integrated into kernel t-SNE based on the Fisher information. (C) 2014 Elsevier B.V. All rights reserved.