Clustering in the Presence of Scatter

Clustering in the Presence of Scatter
复制标题

DOI:
10.1111/j.1541-0420.2008.01064.x
复制
发表时间:
2009-06-01
期刊:
影响因子:
1.9
通讯作者:
Ramler, Ivan P.
Ramler, Ivan P.
中科院分区:
数学3区
文献类型:
--
作者:
Maitra, Ranjan;Ramler, Ivan P.

文献摘要

被引文献

相似文献

提出了一种新的方法来聚类数据集的存在下,分散的观察。分散的观察被定义为与任何其他观察都不同,因此将它们分组的传统方法可能会导致错误的结论。我们建议的方法是一个计划,假设均匀的球形集群,迭代地建立核心周围的中心和组内的每个核心点,同时确定点以外的分散。在没有分散的情况下,该算法简化为k-均值。我们还提供了初始化算法和估计数据集中聚类数的方法。实验结果表明,该算法具有很好的性能,特别是当簇是椭圆对称的时候。该方法被应用于分析美国环境保护署的有毒物质排放清单报告的工业汞排放的2000年。
A new methodology is proposed for clustering datasets in the presence of scattered observations. Scattered observations are defined as unlike any other, so traditional approaches that force them into groups can lead to erroneous conclusions. Our suggested approach is a scheme which, under assumption of homogeneous spherical clusters, iteratively builds cores around their centers and groups points within each core while identifying points outside as scatter. In the absence of scatter, the algorithm reduces to k-means. We also provide methodology to initialize the algorithm and to estimate the number of clusters in the dataset. Results in experimental situations show excellent performance, especially when clusters are elliptically symmetric. The methodology is applied to the analysis of the United States Environmental Protection Agency's Toxic Release Inventory reports on industrial releases of mercury for the year 2000.