Stochastic k-Neighborhood Selection for Supervised and Unsupervised Learning

Stochastic k-Neighborhood Selection for Supervised and Unsupervised Learning
复制标题

DOI:
--
复制
发表时间:
2013-06
期刊:
--
影响因子:
--
通讯作者:
Daniel Tarlow;Kevin Swersky;Laurent Charlin;I. Sutskever;R. Zemel
Daniel Tarlow;Kevin Swersky;Laurent Charlin;I. Sutskever;R. Zemel
中科院分区:
其他
文献类型:
--
作者:
Daniel Tarlow;Kevin Swersky;Laurent Charlin;I. Sutskever;R. Zemel

文献摘要

被引文献

相似文献

邻域成分分析(NCA)是一种流行的方法,用于学习在k-最近邻(kNN)分类器中使用的距离度量。模型中内置的一个关键假设是,每个点随机选择一个邻居,这使得模型仅对k = 1的kNN是合理的。然而,k > 1的kNN分类器更鲁棒,并且在实践中通常是优选的。在这里,我们提出了kNCA,它通过学习适合于任意k的kNN的距离度量来推广NCA。主要的技术贡献是展示了如何有效地计算和优化kNN分类器的预期精度。我们在无监督设置中应用类似的想法来产生kSNE和kt-SNE,这是随机邻居嵌入(SNE,t-SNE)的推广,它在大小为k的邻域上操作,它提供了一个控制嵌入的轴,允许更均匀和可解释的区域。从经验上讲,我们表明,kNCA往往提高了最先进的方法的分类精度,产生定性差异的嵌入k是不同的,是更强大的标签噪声。
Neighborhood Components Analysis (NCA) is a popular method for learning a distance metric to be used within a k-nearest neighbors (kNN) classifier. A key assumption built into the model is that each point stochastically selects a single neighbor, which makes the model well-justified only for kNN with k = 1. However, kNN classifiers with k > 1 are more robust and usually preferred in practice. Here we present kNCA, which generalizes NCA by learning distance metrics that are appropriate for kNN with arbitrary k. The main technical contribution is showing how to efficiently compute and optimize the expected accuracy of a kNN classifier. We apply similar ideas in an unsupervised setting to yield kSNE and kt-SNE, generalizations of Stochastic Neighbor Embedding (SNE, t-SNE) that operate on neighborhoods of size k, which provide an axis of control over embeddings that allow for more homogeneous and interpretable regions. Empirically, we show that kNCA often improves classification accuracy over state of the art methods, produces qualitative differences in the embeddings as k is varied, and is more robust with respect to label noise.