Selection of Variables for Cluster Analysis and Classification Rules

Selection of Variables for Cluster Analysis and Classification Rules
复制标题

DOI:
10.1198/016214508000000544
复制
发表时间:
2008-09-01
影响因子:
3.7
通讯作者:
Svarc, Marcela
Svarc, Marcela
中科院分区:
数学1区
文献类型:
--
作者:
Fraiman, Ricardo;Justel, Ana;Svarc, Marcela

文献摘要

被引文献

相似文献

本文介绍了聚类分析中变量选择和分类规则的两种方法。一种主要是检测“噪声”的非信息性变量,而另一种则也涉及多线性和一般相关性。这两种方法都被设计为在执行了“令人满意的”分组过程之后使用。提出了一种前向-后向算法,使这类过程在大数据集上可行。进行了小型仿真,并对一些实际数据进行了分析。
In this article we introduce two procedures for variable selection in cluster analysis and classification rules. One is mainly aimed at detecting the ''noisy'' noninformative variables, while the other also deals with multicolinearity and general dependence. Both methods are designed to be used after a ''satisfactory'' grouping procedure has been carried out. A forward-backward algorithm is proposed to make such procedures feasible in large datasets. A small simulation is performed and some real data examples are analyzed.