SDE: A Novel Clustering Framework Based on Sparsity-Density Entropy

SDE: A Novel Clustering Framework Based on Sparsity-Density Entropy
复制标题

DOI:
10.1109/tkde.2018.2792021
复制
发表时间:
2018-08
影响因子:
8.9
通讯作者:
Sheng Li;Lusi Li;Jun Yan;Haibo He
Sheng Li;Lusi Li;Jun Yan;Haibo He
中科院分区:
计算机科学2区
文献类型:
--
作者:
Sheng Li;Lusi Li;Jun Yan;Haibo He

文献摘要

被引文献

相似文献

高维变密度数据的聚类对传统的基于密度的聚类方法提出了挑战。近年来,熵作为信息不确定性的一种数值度量,可以用来度量样本在数据空间中的边界程度,也可以用来选择特征集中的重要特征。在我们的新框架中,基于稀疏密度熵(sparsity-density entropy,简称NTD)对高维、变密度的数据进行聚类。该算法首先对多维数据进行高质量采样,并利用稀疏分数熵(SSE)选择具有代表性的特征。然后采用一种新的变密度聚类方法--密度熵(DE)方法,得到聚类结果和噪声。DE自动确定边界集的基础上的边界度的全局最小值,然后自适应地执行聚类分析的基础上的每个局部聚类的边界度的局部最小值。在合成数据集和真实的数据集上,通过与几种聚类算法的比较,验证了该框架的有效性和效率。结果表明,该框架能够同时检测噪声和处理高维、不同密度的数据。
Clustering of data with high dimension and variable densities poses a remarkable challenge to the traditional density-based clustering methods. Recently, entropy, a numerical measure of the uncertainty of information, can be used to measure the border degree of samples in data space and also select significant features in feature set. It was used in our new framework based on the sparsity-density entropy (SDE) to cluster the data with high dimension and variable densities. First, SDE conducts high-quality sampling for multidimensional data and selects the representative features using sparsity score entropy (SSE). Second, the clustering results and noises are obtained adopting a new density-variable clustering method called density entropy (DE). DE automatically determines the border set based on the global minimum of border degrees and then adaptively performs cluster analysis for each local cluster based on the local minimum of border degrees. The effectiveness and efficiency of the proposed SDE framework are validated on synthetic and real data sets in comparison with several clustering algorithms. The results showed that the proposed SDE framework concurrently detected the noises and processed the data with high dimension and various densities.