A criterion for variable selection in multiple discriminant analysis

A criterion for variable selection in multiple discriminant analysis
复制标题

多重判别分析中变量选择的标准

DOI:
10.32917/hmj/1206133544
复制
发表时间:
1983
影响因子:
0.2
通讯作者:
Y. Fujikoshi
Y. Fujikoshi
中科院分区:
数学4区
文献类型:
--
作者:
Y. Fujikoshi

文献摘要

被引文献

相似文献

本文讨论了q + 1总体多元线性模型判别分析中的变量选择问题。变量的选择很重要,因为在某些情况下,从原始变量中删除某些变量可能更适合统计分析的实际目的。已经提出了许多逐步的程序来减少在q + 1个群体之间进行区分所需的变量的数量(例如,参见McCabe [7]、Farmer和Freund [2])。McKay [8]提出了一个程序,用于确定所有的变量子集,这些变量子集在q +1总体之间提供的分离度基本上与原始变量集一样多,基于Gabriel [6]意义上的同时检验程序。本文提出了一个在判别分析中确定“最佳”变量子集的标准,其目的是仅根据几个典型判别变量来解释q + 1总体之间的差异。我们得到一个credibility,基于模型拟合的方法。通过引入参数模型族,我们把寻找变量的“最佳”子集的问题看作是寻找“最佳”模型的问题。参数模型基于Rao的“无附加信息假设”[10]。我们的准则是通过应用Akaike的信息准则(Akaike [1])来选择模型而得到的。本文还讨论了在多元线性模型中寻找变量的“最佳”子集的问题。这是判别分析中变量选择问题的推广。得到了多元线性模型中变量选择准则的渐近分布,推广了Fujikoshi [5]在两组判别分析中的结果.在原始变量是先验有序的情况下,渐近分布可以简化为一个简单的形式。
This paper deals with the problem of variable selection in discriminant analysis with q + 1 populations and a multivariate linear model. The variable selection is important since there are situations where the deletion of some variables from the original variables may be preferable for the practical aim of statistical analysis. A number of step wise procedures have been proposed for reducing the number of variables required to discriminate among the q + 1 populations (e.g., see McCabe [7], Farmer and Freund [2]). McKay [8] has proposed a procedure for determining all subsets of variables that provide essentially as much separation among the q +1 populations as the original set of variables, based on a simultaneous test procedure in Gabriel's [6] sense. In this paper we propose a criterion for determining the "best" subset of variables in the discriminant analysis whose aim is to interpret the differences among the q + 1 populations in terms of only a few canonical discriminant variables. We obtain a creterion, based on a model fitting approach. We regard the problem of finding the "best" subset of variables as one of finding the "best" model, by introducing a family of parametric models. The parametric models are based on "no additional information hypotheses" due to Rao [10]. Our criterion is obtained by applying Akaike's information criterion (Akaike [1]) to choice of the models. The problem of finding the "best" subset of variables in a multivariate linear model is also discussed. This is a generalization of the problem of variable selection in the discriminant analysis. Asymptotic distributions of the criterion for variable selection in the multivariate linear model are obtained, resulting in generalizations of Fujikoshi [5] in the case of two-group discriminant analysis. The asymptotic distribution in the case when the original variables are ordered a priori can be reduced to a simple form.