Simulation-based hypothesis testing of high dimensional means under covariance heterogeneity

Simulation-based hypothesis testing of high dimensional means under covariance heterogeneity
复制标题

协方差异质性下基于模拟的高维均值假设检验

DOI:
10.1111/biom.12695
复制
发表时间:
2017-12-01
期刊:
影响因子:
1.9
通讯作者:
Zhou, Wen
Zhou, Wen
中科院分区:
数学3区
文献类型:
--
作者:
Chang, Jinyuan;Zheng, Chao;Zhou, Wen

文献摘要

被引文献

相似文献

本文研究了高维数据在单样本和双样本情况下均值向量的检验问题。建议的测试程序采用最大型统计和参数引导技术来计算临界值。与现有的检验严重依赖于未知协方差矩阵的结构性条件不同,本文提出的检验允许数据具有一般的协方差结构,因而具有广泛的实用性。为了提高权力的测试对稀疏的替代品,我们进一步提出了两个步骤的程序与初步的功能筛选步骤。所提出的测试的理论特性进行了研究。通过对合成数据集和人类急性淋巴细胞白血病基因表达数据集进行广泛的数值实验,我们说明了新测试的性能以及它们如何为检测疾病相关基因集提供帮助。所提出的方法已在R-包HDtest中实现,并可在CRAN上使用。
In this article, we study the problem of testing the mean vectors of high dimensional data in both one-sample and two-sample cases. The proposed testing procedures employ maximum-type statistics and the parametric bootstrap techniques to compute the critical values. Different from the existing tests that heavily rely on the structural conditions on the unknown covariance matrices, the proposed tests allow general covariance structures of the data and therefore enjoy wide scope of applicability in practice. To enhance powers of the tests against sparse alternatives, we further propose two-step procedures with a preliminary feature screening step. Theoretical properties of the proposed tests are investigated. Through extensive numerical experiments on synthetic data sets and an human acute lymphoblastic leukemia gene expression data set, we illustrate the performance of the new tests and how they may provide assistance on detecting disease-associated gene-sets. The proposed methods have been implemented in an R-package HDtest and are available on CRAN.