An Algorithm for Clustering Categorical Data With Set-Valued Features

An Algorithm for Clustering Categorical Data With Set-Valued Features
复制标题

具有集值特征的分类数据聚类算法

DOI:
10.1109/tnnls.2017.2770167
复制
发表时间:
2018-10
影响因子:
10.4
通讯作者:
Qian Yuhua
Qian Yuhua
中科院分区:
计算机科学1区
文献类型:
--
作者:
Cao Fuyuan;Huang Joshua Zhexue;Liang Jiye;Zhao Xingwang;Meng Yinfeng;Feng Kai;Qian Yuhua

文献摘要

参考文献

被引文献

相似文献

在数据挖掘中,对象通常由一组特征表示,其中对象的每个特征只有一个值。然而,在现实中,一些特征可以具有多个值,例如,一个人有多个职称、爱好和电子邮件地址。这些特征可以被称为集值特征,并且在使用现有的数据挖掘算法来分析具有集值特征的数据时,这些特征通常被用伪特征来处理。在本文中,我们提出了一个SV-<inline-formula><tex-math notation="LaTeX">$k$</tex-math></inline-formula>-模式算法,聚类与集值特征的分类数据。在该算法中,定义了两个具有集值特征的对象之间的距离函数,并提出了聚类中心的集值模式表示。我们开发了一种启发式的方法来更新聚类中心的迭代聚类过程和初始化算法来选择初始聚类中心。分析了SV-<inline-formula><tex-math notation="LaTeX">$k$</tex-math></inline-formula>-modes算法的收敛性和复杂度.实验进行了合成数据和真实的数据从五个不同的应用。实验结果表明,SV-<inline-formula><tex-math notation="LaTeX">$k$</tex-math></inline-formula>-modes算法在聚类真实的数据时比其他三种分类聚类算法性能更好,并且该算法对大数据具有可扩展性。
In data mining, objects are often represented by a set of features, where each feature of an object has only one value. However, in reality, some features can take on multiple values, for instance, a person with several job titles, hobbies, and email addresses. These features can be referred to as set-valued features and are often treated with dummy features when using existing data mining algorithms to analyze data with set-valued features. In this paper, we propose an SV-<inline-formula> <tex-math notation="LaTeX">$k$ </tex-math></inline-formula>-modes algorithm that clusters categorical data with set-valued features. In this algorithm, a distance function is defined between two objects with set-valued features, and a set-valued mode representation of cluster centers is proposed. We develop a heuristic method to update cluster centers in the iterative clustering process and an initialization algorithm to select the initial cluster centers. The convergence and complexity of the SV-<inline-formula> <tex-math notation="LaTeX">$k$ </tex-math></inline-formula>-modes algorithm are analyzed. Experiments are conducted on both synthetic data and real data from five different applications. The experimental results have shown that the SV-<inline-formula> <tex-math notation="LaTeX">$k$ </tex-math></inline-formula>-modes algorithm performs better when clustering real data than do three other categorical clustering algorithms and that the algorithm is scalable to large data.
在不断变化的数据流上进行基于增量密度的集成聚类
DOI: 10.1016/j.neucom.2016.01.009
发表时间: 2016-05
期刊: Neurocomputing
影响因子: 6
作者:
Imran Khan;Joshua Zhexue Huang;Kamen Ivanov
通讯作者: Kamen Ivanov
DOI: 10.1007/s10489-007-0111-x
发表时间: 2009-08
影响因子: 5.3
作者:
Min-Ling Zhang;Zhi-Hua Zhou
通讯作者: Min-Ling Zhang;Zhi-Hua Zhou
DOI: --
发表时间: 2001
期刊: --
影响因子: --
作者:
通讯作者: --
DOI: 10.1016/j.patcog.2015.05.006
发表时间: 2015-11-01
影响因子: 8
作者:
Jing, Liping;Tian, Kuang;Huang, Joshua Z.
通讯作者: Huang, Joshua Z.
DOI: 10.1038/234034a0
发表时间: 1971-01-01
期刊: NATURE
影响因子: 64.8
作者:
LEVANDOWSKY, M;WINTER, D
通讯作者: WINTER, D