An Experimental Comparison of GPU Techniques for DBSCAN Clustering

An Experimental Comparison of GPU Techniques for DBSCAN Clustering
复制标题

DBSCAN 聚类 GPU 技术的实验比较

DOI:
--
复制
发表时间:
2019
期刊:
2019 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
L. Gruenwald
L. Gruenwald
中科院分区:
--
文献类型:
--
作者:
Hamza Mustafa;Eleazar Leal;L. Gruenwald

文献摘要

被引文献

相似文献

DBSCAN是一种基于密度的群集算法,特别适用于查找任意形状的群集。与K-Means等其他聚类技术不同,它不需要将聚类的数量指定为输入参数,并且对孤立点具有高度的健壮性。然而,DBSCAN具有最坏情况下的二次型时间复杂性,这使得处理大型数据集变得困难。为了解决这个问题,已经提出了几项工作来利用DBSCAN集群中GPU的巨大并行性。尽管如此,这些作品都没有被实验地相互比较。在本文中,我们回顾了现有的用于DBSCAN聚类的GPU算法,并使用三个真实数据集对这些GPU算法进行了首次实验研究,以确定性能最好的算法。实验结果表明,在执行时间和内存需求方面,CUDA-DCLUST是性能最好的GPU算法。
DBSCAN is a density-based clustering algorithm that is especially useful for finding clusters of arbitrary shapes. As opposed to other clustering techniques, like K-means, it does not require the number of clusters to be specified as an input parameter, and it is highly robust to outliers. However, DBSCAN has a worst-case quadratic time complexity, which makes it difficult to handle large dataset sizes. To address this problem, several works have been proposed that exploit the massive parallelism of GPUs in DBSCAN clustering. Nonetheless, none of these works have been experimentally compared against each other. In this paper, we review the existing GPU algorithms for DBSCAN clustering and conduct the first experimental study comparing these GPU algorithms using three real-world datasets to identify the best performing algorithm. Our results show that CUDA-DClust is the best performing GPU algorithm in terms of execution time and memory requirements.