Mp-Dissimilarity: A Data Dependent Dissimilarity Measure

Mp-Dissimilarity: A Data Dependent Dissimilarity Measure
复制标题

DOI:
10.1109/icdm.2014.33
复制
发表时间:
2014-12
期刊:
2014 IEEE International Conference on Data Mining
影响因子:
--
通讯作者:
Sunil Aryal;K. Ting;Gholamreza Haffari;T. Washio
Sunil Aryal;K. Ting;Gholamreza Haffari;T. Washio
中科院分区:
其他
文献类型:
--
作者:
Sunil Aryal;K. Ting;Gholamreza Haffari;T. Washio

文献摘要

被引文献

相似文献

最近邻搜索是许多数据挖掘算法的核心过程。在高维空间中找到查询的可靠最接近匹配仍然是一项具有挑战性的任务。这是因为许多基于几何模型的相异性度量的有效性,例如lp范数,随着维数的增加而降低。在本文中,我们研究了如何利用数据分布来衡量两个实例之间的相异性,并提出了一个新的数据相关的相异性度量称为“MP-差异”。它不依赖于几何距离,而是将每个维度中两个实例之间的相异性度量为包围两个实例的区域中的概率质量。它认为稀疏区域中的两个实例比密集区域中的两个实例更相似,尽管这两对实例具有相同的几何距离。我们的实证结果表明,所提出的相异性测度确实提供了一个可靠的最近邻搜索在高维空间,特别是在稀疏数据。在分类和信息检索任务中,mp-不相似性比lp-范数和余弦距离产生更好的任务特定性能。
Nearest neighbour search is a core process in many data mining algorithms. Finding reliable closest matches of a query in a high dimensional space is still a challenging task. This is because the effectiveness of many dissimilarity measures, that are based on a geometric model, such as lp-norm, decreases as the number of dimensions increases. In this paper, we examine how the data distribution can be exploited to measure dissimilarity between two instances and propose a new data dependent dissimilarity measure called 'mp-dissimilarity'. Rather than relying on geometric distance, it measures the dissimilarity between two instances in each dimension as a probability mass in a region that encloses the two instances. It deems the two instances in a sparse region to be more similar than two instances in a dense region, though these two pairs of instances have the same geometric distance. Our empirical results show that the proposed dissimilarity measure indeed provides a reliable nearest neighbour search in high dimensional spaces, particularly in sparse data. Mp-dissimilarity produced better task specific performance than lp-norm and cosine distance in classification and information retrieval tasks.