Parallel Canopy Clustering on GPUs
Parallel Canopy Clustering on GPUs
复制标题
DOI:
10.1007/978-3-319-22849-5_23
复制
发表时间:
2015-09
期刊:
影响因子:
--
通讯作者:
Yusuke Kozawa;Fumitaka Hayashi;Toshiyuki Amagasa;H. Kitagawa
中科院分区:
文献类型:
--
作者:
Yusuke Kozawa;Fumitaka Hayashi;Toshiyuki Amagasa;H. Kitagawa
Canopy clustering is a preprocessing method for standard clustering algorithms such ask-means and hierarchical agglomerative clustering. Canopy clustering can greatly reduce the computational cost of clustering algorithms. However, canopy clustering itself may also take a vast amount of time for handling massive data, if we naïvely implement it. To address this problem, we present efficient algorithms and implementations of canopy clustering on GPUs, which have evolved recently as general-purpose many-core processors. We not only accelerate the computation of original canopy clustering, but also propose an algorithm using grid index. This algorithm partitions the data into cells to reduce redundant computations and, at the same time, to exploit the parallelism of GPUs. Experiments show that the proposed implementations on the GPU is 2 times faster on average than multi-threaded, SIMD implementations on two octa-core CPUs.