Two-Stage Procedures for High-Dimensional Data

Two-Stage Procedures for High-Dimensional Data
复制标题

DOI:
10.1080/07474946.2011.619088
复制
发表时间:
2011-01-01
影响因子:
0.8
通讯作者:
Yata, Kazuyoshi
Yata, Kazuyoshi
中科院分区:
数学4区
文献类型:
--
作者:
Aoshima, Makoto;Yata, Kazuyoshi

文献摘要

被引文献

相似文献

在这篇文章中,我们考虑了高维数据的各种推理问题。本文的目的是提出未来的研究方向和可能的解决方案的p >> n问题,使用新类型的两阶段估计方法。这是首次尝试将序贯分析应用于高维统计推断,以确保预先指定的准确性。我们通过创建新类型的多变量两阶段程序来提供推理问题的样本量确定。为了发展理论和方法,最重要和基本的思想是p ->无穷大时的渐进正态性。通过发展当p ->无穷大时的渐近正态性,我们首先给出(a)平方损失的给定带宽置信区域。此外,我们给出了(B)一个双样本检验,以确保预先设定的大小和权力,同时与(c)一个平等的检验程序的两个协方差矩阵。我们还给出了(d)一个两阶段的判别过程,控制错误分类率不超过一个预先设定的值。此外,我们建议(e)一个两阶段的变量选择程序,提供筛选的变量在第一阶段,并选择一个重要的一组相关的变量从一组候选变量在第二阶段。在变量选择过程之后,我们考虑(f)高维回归的变量选择,以便在精度保证和计算成本方面与套索进行比较。此外,我们考虑变量选择分类,并提出(g)一个两阶段的判别程序后,筛选一些变量。最后,我们考虑(h)通过构建相关系数的多重检验来进行高维数据的通径分析。
In this article, we consider a variety of inference problems for high-dimensional data. The purpose of this article is to suggest directions for future research and possible solutions about p >> n problems by using new types of two-stage estimation methodologies. This is the first attempt to apply sequential analysis to high-dimensional statistical inference ensuring prespecified accuracy. We offer the sample size determination for inference problems by creating new types of multivariate two-stage procedures. To develop theory and methodologies, the most important and basic idea is the asymptotic normality when p -> infinity. By developing asymptotic normality when p -> infinity, we first give (a) a given-bandwidth confidence region for the square loss. In addition, we give (b) a two-sample test to assure prespecified size and power simultaneously together with (c) an equality-test procedure for two covariance matrices. We also give (d) a two-stage discriminant procedure that controls misclassification rates being no more than a prespecified value. Moreover, we propose (e) a two-stage variable selection procedure that provides screening of variables in the first stage and selects a significant set of associated variables from among a set of candidate variables in the second stage. Following the variable selection procedure, we consider (f) variable selection for high-dimensional regression to compare favorably with the lasso in terms of the assurance of accuracy and the computational cost. Further, we consider variable selection for classification and propose (g) a two-stage discriminant procedure after screening some variables. Finally, we consider (h) pathway analysis for high-dimensional data by constructing a multiple test of correlation coefficients.