A criterion for variable selection in multiple discriminant analysis
A criterion for variable selection in multiple discriminant analysis
复制标题
多重判别分析中变量选择的标准
DOI:
10.32917/hmj/1206133544
复制
发表时间:
1983
影响因子:
0.2
通讯作者:
Y. Fujikoshi
中科院分区:
文献类型:
--
作者:
Y. Fujikoshi
This paper deals with the problem of variable selection in discriminant analysis with q + 1 populations and a multivariate linear model. The variable selection is important since there are situations where the deletion of some variables from the original variables may be preferable for the practical aim of statistical analysis. A number of step wise procedures have been proposed for reducing the number of variables required to discriminate among the q + 1 populations (e.g., see McCabe [7], Farmer and Freund [2]). McKay [8] has proposed a procedure for determining all subsets of variables that provide essentially as much separation among the q +1 populations as the original set of variables, based on a simultaneous test procedure in Gabriel's [6] sense. In this paper we propose a criterion for determining the "best" subset of variables in the discriminant analysis whose aim is to interpret the differences among the q + 1 populations in terms of only a few canonical discriminant variables. We obtain a creterion, based on a model fitting approach. We regard the problem of finding the "best" subset of variables as one of finding the "best" model, by introducing a family of parametric models. The parametric models are based on "no additional information hypotheses" due to Rao [10]. Our criterion is obtained by applying Akaike's information criterion (Akaike [1]) to choice of the models. The problem of finding the "best" subset of variables in a multivariate linear model is also discussed. This is a generalization of the problem of variable selection in the discriminant analysis. Asymptotic distributions of the criterion for variable selection in the multivariate linear model are obtained, resulting in generalizations of Fujikoshi [5] in the case of two-group discriminant analysis. The asymptotic distribution in the case when the original variables are ordered a priori can be reduced to a simple form.