Nonsmooth Penalized Clustering via Regularized Sparse Regression

Nonsmooth Penalized Clustering via Regularized Sparse Regression
复制标题

通过正则化稀疏回归的非平滑惩罚聚类

DOI:
10.1109/tcyb.2016.2546965
复制
发表时间:
2016
影响因子:
11.8
通讯作者:
Zhiquan Qi
Zhiquan Qi
中科院分区:
计算机科学1区
文献类型:
--
作者:
Lingfeng Niu;Ruizhi Zhou;Yingjie Tian;Zhiquan Qi

文献摘要

被引文献

相似文献

聚类在数据分析中得到了广泛的应用。大多数现有的聚类方法都假设聚类的数目是预先给定的。最近,提出了一种新的聚类框架,它可以从训练数据中自动学习聚类的数量。在此基础上,我们提出了一种基于正则稀疏回归的非光滑惩罚聚类模型。特别是,该模型被制定为一个非光滑的非凸优化,这是基于过参数化,并利用基于范数的正则化来控制模型拟合和集群的数量之间的权衡。从理论上证明了新模型能够保证聚类中心的稀疏性。为了增加其实用性,实际使用中,我们坚持一个易于计算的标准,并遵循一个策略,以缩小搜索区间的交叉验证。针对代价函数的非光滑性和非凸性,提出了一种简单的光滑信赖域算法,并给出了算法的收敛性和计算复杂度分析。模拟和实际数据集的数值研究提供了支持,我们的理论结果,并证明了我们的新方法的优点。
Clustering has been widely used in data analysis. A majority of existing clustering approaches assume that the number of clusters is given in advance. Recently, a novel clustering framework is proposed which can automatically learn the number of clusters from training data. Based on these works, we propose a nonsmooth penalized clustering model via() regularized sparse regression. In particular, this model is formulated as a nonsmooth nonconvex optimization, which is based on over-parameterization and utilizes an-norm-based regularization to control the tradeoff between the model fit and the number of clusters. We theoretically prove that the new model can guarantee the sparseness of cluster centers. To increase its practicality for practical use, we adhere to an easy-to-compute criterion and follow a strategy to narrow down the search interval of cross validation. To address the nonsmoothness and nonconvexness of the cost function, we propose a simple smoothing trust region algorithm and present its convergent and computational complexity analysis. Numerical studies on both simulated and practical data sets provide support to our theoretical results and demonstrate the advantages of our new method.