Survey of State-of-the-Art Mixed Data Clustering Algorithms

Survey of State-of-the-Art Mixed Data Clustering Algorithms
复制标题

DOI:
10.1109/access.2019.2903568
复制
发表时间:
2019-01-01
期刊:
影响因子:
3.9
通讯作者:
Khan, Shehroz S.
Khan, Shehroz S.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ahmad, Amir;Khan, Shehroz S.

文献摘要

被引文献

相似文献

混合数据包括数字和分类特征,混合数据集经常出现在许多领域,如健康,金融和营销。聚类通常应用于混合数据集,以查找结构并将相似对象分组以供进一步分析。然而,对混合数据进行聚类具有挑战性,因为很难将求和或平均等数学运算直接应用于这些数据集的特征值。在本文中,我们提出了一个分类的混合数据聚类算法的研究,确定五个主要的研究主题。然后,我们提出了国家的最先进的审查的研究工作在每个研究主题。我们分析了这些方法的优点和缺点,并指出了未来的研究方向。最后,我们提出了在这一领域的整体挑战的深入分析,突出开放的研究问题,并讨论在该领域取得进展的指导方针。
Mixed data comprises both numeric and categorical features, and mixed datasets occur frequently in many domains, such as health, finance, and marketing. Clustering is often applied to mixed datasets to find structures and to group similar objects for further analysis. However, clustering mixed data are challenging because it is difficult to directly apply mathematical operations, such as summation or averaging, to the feature values of these datasets. In this paper, we present a taxonomy for the study of mixed data clustering algorithms by identifying five major research themes. We then present the state-of-the-art review of the research works within each research theme. We analyze the strengths and weaknesses of these methods with pointers for future research directions. At last, we present an in-depth analysis of the overall challenges in this field, highlight open research questions, and discuss guidelines to make progress in the field.