Simple and Scalable Sparse k-means Clustering via Feature Ranking

Simple and Scalable Sparse k-means Clustering via Feature Ranking
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
arXiv: Machine Learning
影响因子:
--
通讯作者:
Zhiyue Zhang;K. Lange;Jason Xu
Zhiyue Zhang;K. Lange;Jason Xu
中科院分区:
其他
文献类型:
--
作者:
Zhiyue Zhang;K. Lange;Jason Xu

文献摘要

相似文献

聚类是无监督学习中的一项基本活动,当特征空间是高维时,聚类是非常困难的。幸运的是,在许多现实场景中,只有少数特征与区分聚类相关。这激发了稀疏聚类技术的发展,这种技术通常依赖于高计算复杂度的外部算法中的k均值。目前的技术还需要仔细调整收缩参数,进一步限制了它们的可扩展性。在本文中,我们提出了一个新的框架,稀疏k均值聚类,是直观的,简单的实现,并具有竞争力的国家的最先进的算法。我们表明,我们的算法享有一致性和收敛性的保证。我们的核心方法很容易推广到几个特定于任务的算法,如聚类属性的子集和部分观察到的数据设置。我们通过模拟实验和真实的数据基准,包括三体小鼠蛋白质表达的案例研究,彻底展示了这些贡献。
Clustering, a fundamental activity in unsupervised learning, is notoriously difficult when the feature space is high-dimensional. Fortunately, in many realistic scenarios, only a handful of features are relevant in distinguishing clusters. This has motivated the development of sparse clustering techniques that typically rely on k-means within outer algorithms of high computational complexity. Current techniques also require careful tuning of shrinkage parameters, further limiting their scalability. In this paper, we propose a novel framework for sparse k-means clustering that is intuitive, simple to implement, and competitive with state-of-the-art algorithms. We show that our algorithm enjoys consistency and convergence guarantees. Our core method readily generalizes to several task-specific algorithms such as clustering on subsets of attributes and in partially observed data settings. We showcase these contributions thoroughly via simulated experiments and real data benchmarks, including a case study on protein expression in trisomic mice.