Variable selection for model-based clustering using the integrated complete-data likelihood

Variable selection for model-based clustering using the integrated complete-data likelihood
复制标题

DOI:
10.1007/s11222-016-9670-1
复制
发表时间:
2017-07-01
影响因子:
2.2
通讯作者:
Sedki, Mohammed
Sedki, Mohammed
中科院分区:
数学2区
文献类型:
--
作者:
Marbac, Matthieu;Sedki, Mohammed

文献摘要

被引文献

相似文献

聚类分析中的变量选择是一个重要而又具有挑战性的问题。它可以通过正则化方法来实现,该方法通过使用拉索型惩罚来实现聚类精度和所选变量数量之间的权衡。然而,惩罚项的校准可能会受到批评。模型选择方法是一种有效的替代方法,但它们需要一个困难的信息标准,涉及组合问题的优化。首先,大多数这些优化算法是基于一个次优的过程(例如逐步的方法)。其次,算法通常是计算昂贵的,因为它们需要多次调用EM算法。在这里,我们建议使用一个新的信息标准的基础上集成的完整数据的可能性。它不需要最大似然估计和它的最大化似乎是简单和计算效率。我们的方法的原始贡献是执行模型选择,而不需要任何参数估计。然后,仅需要对唯一选择的模型进行参数推断。该方法用于条件独立的高斯混合模型的变量选择。在模拟数据集和基准数据集上的数值实验表明,该方法的性能通常优于两种经典的变量选择方法。所提出的方法在CRAN上可用的R包VarSelLCM中实现。
Variable selection in cluster analysis is important yet challenging. It can be achieved by regularization methods, which realize a trade-off between the clustering accuracy and the number of selected variables by using a lasso-type penalty. However, the calibration of the penalty term can suffer from criticisms. Model selection methods are an efficient alternative, yet they require a difficult optimization of an information criterion which involves combinatorial problems. First, most of these optimization algorithms are based on a suboptimal procedure (e.g. stepwise method). Second, the algorithms are often computationally expensive because they need multiple calls of EM algorithms. Here we propose to use a new information criterion based on the integrated complete-data likelihood. It does not require the maximum likelihood estimate and its maximization appears to be simple and computationally efficient. The original contribution of our approach is to perform the model selection without requiring any parameter estimation. Then, parameter inference is needed only for the unique selected model. This approach is used for the variable selection of a Gaussian mixture model with conditional independence assumed. The numerical experiments on simulated and benchmark datasets show that the proposed method often outperforms two classical approaches for variable selection. The proposed approach is implemented in the R package VarSelLCM available on CRAN.