Choosing lp norms in high-dimensional spaces based on hub analysis

Choosing lp norms in high-dimensional spaces based on hub analysis
复制标题

DOI:
10.1016/j.neucom.2014.11.084
复制
发表时间:
2015-12-02
期刊:
影响因子:
6
通讯作者:
Schnitzer, Dominik
Schnitzer, Dominik
中科院分区:
计算机科学2区
文献类型:
--
作者:
Flexer, Arthur;Schnitzer, Dominik

文献摘要

被引文献

相似文献

中心现象是最近发现的维数诅咒的一个方面。中心对象与大量数据点的距离较小,而反中心对象距离所有其他数据点较远。一个密切相关的问题是高维空间中距离的集中。之前的工作已经提倡使用分数 l(p) 范数而不是普遍存在的欧几里得范数,以避免距离集中的负面影响。然而,使用哪种精确的分数范数在很大程度上是一个尚未解决的问题。这项工作的贡献是对不同 l(P) 规范和中心性之间的关系进行了实证分析。我们提出了一种无监督方法来选择 l(P) 范数,该范数可以最小化集线器,同时最大化最近邻分类。我们的方法在七个高维数据集上进行了评估,并与重新缩放距离以避免中心化的三种方法进行了比较。 (C) 2015 年作者。由 Elsevier B.V 出版。这是一篇基于 CC BY 许可证 (http://creativecommons.org/licenses/by/4.0/) 的开放获取文章。
The hubness phenomenon is a recently discovered aspect of the curse of dimensionality. Hub objects have a small distance to an exceptionally large number of data points while anti-hubs lie far from all other data points. A closely related problem is the concentration of distances in high-dimensional spaces. Previous work has already advocated the use of fractional l(p) norms instead of the ubiquitous Euclidean norm to avoid the negative effects of distance concentration. However, which exact fractional norm to use is a largely unsolved problem. The contribution of this work is an empirical analysis of the relation of different l(P) norms and hubness. We propose an unsupervised approach for choosing an l(P) norm which minimizes hubs while simultaneously maximizing nearest neighbor classification. Our approach is evaluated on seven high-dimensional data sets and compared to three approaches that re-scale distances to avoid hubness. (C) 2015 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).