Rough-DBSCAN: A fast hybrid density based clustering method for large data sets

Rough-DBSCAN: A fast hybrid density based clustering method for large data sets
复制标题

DOI:
10.1016/j.patrec.2009.08.008
复制
发表时间:
2009-12
期刊:
Pattern Recognit. Lett.
影响因子:
--
通讯作者:
P. Viswanath;V. S. Babu
P. Viswanath;V. S. Babu
中科院分区:
其他
文献类型:
--
作者:
P. Viswanath;V. S. Babu

文献摘要

被引文献

相似文献

基于密度的聚类技术(例如 DBSCAN)很有吸引力,因为它可以找到任意形状的聚类以及噪声异常值。它的时间要求是 O(n2),其中 n 是数据集的大小,因此它不适合处理大型数据集。本文提出的解决方案是首先应用领导者聚类方法从数据集中派生出称为领导者的原型,该原型与原型一起也保留了密度信息,然后使用这些领导者来派生基于密度的集群。所提出的称为粗糙 DBSCAN 的混合聚类技术的时间复杂度仅为 O(n),并使用粗糙集理论进行分析。使用合成数据集和真实世界数据集进行实验研究,以将粗糙 DBSCAN 与 DBSCAN 进行比较。结果表明,对于大型数据集,rough-DBSCAN 可以找到与 DBSCAN 类似的聚类,但始终比 DBSCAN 更快。领导者的一些属性也正式确立为原型。
Density based clustering techniques like DBSCAN are attractive because it can find arbitrary shaped clusters along with noisy outliers. Its time requirement is O(n2) where n is the size of the dataset, and because of this it is not a suitable one to work with large datasets. A solution proposed in the paper is to apply the leaders clustering method first to derive the prototypes called leaders from the dataset which along with prototypes preserves the density information also, then to use these leaders to derive the density based clusters. The proposed hybrid clustering technique called rough-DBSCAN has a time complexity of O(n) only and is analyzed using rough set theory. Experimental studies are done using both synthetic and real world datasets to compare rough-DBSCAN with DBSCAN. It is shown that for large datasets rough-DBSCAN can find a similar clustering as found by the DBSCAN, but is consistently faster than DBSCAN. Also some properties of the leaders as prototypes are formally established.