k Nearest Neighbor Similarity Join Algorithm on High-Dimensional Data Using Novel Partitioning Strategy

k Nearest Neighbor Similarity Join Algorithm on High-Dimensional Data Using Novel Partitioning Strategy
复制标题

DOI:
10.1155/2022/1249393
复制
发表时间:
2022-04
影响因子:
--
通讯作者:
Youzhong Ma;Qiaozhi Hua;Zheng Wen;Ruiling Zhang;Yongxin Zhang;Haipeng Li
Youzhong Ma;Qiaozhi Hua;Zheng Wen;Ruiling Zhang;Yongxin Zhang;Haipeng Li
中科院分区:
计算机科学4区
文献类型:
--
作者:
Youzhong Ma;Qiaozhi Hua;Zheng Wen;Ruiling Zhang;Yongxin Zhang;Haipeng Li

文献摘要

相似文献

高维数据的k近邻相似连接在许多领域有着广泛的应用,但仍存在“维数灾难”和数据集规模大等关键问题。首先利用随机投影技术提出了一种新的降维方案,然后设计了两种新的划分策略,包括等宽划分策略和基于距离分裂树的划分策略,最后在上述划分策略的基础上提出了高维数据的k近邻连接算法。对所提方法的性能进行了全面的实验测试,实验结果表明所提方法具有良好的有效性和性能。
k nearest neighbor similarity join on high-dimensional data has broad applications in many fields; several key challenges still exist for this task such as “curse of dimensionality” and large scale of the dataset. A new dimensionality reduction scheme is proposed by using random projection technique, then we design two novel partition strategies, including equal width partition strategy and distance split tree-based partition strategy, and finally, we propose k nearest neighbor join algorithm on high-dimensional data based on the above partition strategies. We conduct comprehensive experiments to test the performance of the proposed approaches, and the experimental results show that the proposed methods have good effectiveness and performance.