A semiparametric method for clustering mixed data

A semiparametric method for clustering mixed data
复制标题

DOI:
10.1007/s10994-016-5575-7
复制
发表时间:
2016-12-01
期刊:
影响因子:
7.5
通讯作者:
Heching, Aliza
Heching, Aliza
中科院分区:
计算机科学3区
文献类型:
--
作者:
Foss, Alex;Markatou, Marianthi;Heching, Aliza

文献摘要

被引文献

相似文献

尽管存在大量的聚类算法,但聚类仍然是一个具有挑战性的问题。随着大型数据集在许多不同领域中变得越来越普遍,通常情况下,聚类算法必须应用于异构变量集,从而迫切需要针对混合连续和分类规模数据的稳健且可扩展的聚类方法。我们表明,如果没有强大的参数假设,当前混合类型数据的聚类方法通常无法公平地平衡连续变量和分类变量的贡献。我们开发了 KAMILA(混合大型数据的 KAy 均值),这是一种直接解决这一基本问题的聚类方法。我们研究了我们的方法的理论方面,并在一系列蒙特卡罗模拟研究和一组实际应用中证明了其有效性。
Despite the existence of a large number of clustering algorithms, clustering remains a challenging problem. As large datasets become increasingly common in a number of different domains, it is often the case that clustering algorithms must be applied to heterogeneous sets of variables, creating an acute need for robust and scalable clustering methods for mixed continuous and categorical scale data. We show that current clustering methods for mixed-type data are generally unable to equitably balance the contribution of continuous and categorical variables without strong parametric assumptions. We develop KAMILA (KAy-means for MIxed LArge data), a clustering method that addresses this fundamental problem directly. We study theoretical aspects of our method and demonstrate its effectiveness in a series of Monte Carlo simulation studies and a set of real-world applications.