Large-Scale Inference of Multivariate Regression for Heavy-Tailed and Asymmetric Data

Large-Scale Inference of Multivariate Regression for Heavy-Tailed and Asymmetric Data
复制标题

DOI:
10.5705/ss.202021.0003
复制
发表时间:
2021
期刊:
影响因子:
1.4
通讯作者:
Youngseok Song;Wen-Xin Zhou;Wen-Xin Zhou
Youngseok Song;Wen-Xin Zhou;Wen-Xin Zhou
中科院分区:
数学3区
文献类型:
--
作者:
Youngseok Song;Wen-Xin Zhou;Wen-Xin Zhou

文献摘要

相似文献

大规模多元回归是一种基本的统计工具,在广泛的领域中得到应用。本文考虑的问题,同时测试了大量的一般线性假设,包括协变量效应分析,方差分析和模型比较。沿着而来的新的挑战是大量的测试是无处不在的存在重尾和/或高度偏斜的测量噪声,这是传统的基于最小二乘的方法失败的主要原因。对于大规模的多元回归,我们开发了一套强大的推理方法来探索数据的功能,如重尾和偏度,这是不可见的最小二乘法的范围。新的测试程序是建立在数据自适应Huber回归,和一个新的协方差估计回归估计。在温和的条件下,我们表明,我们的方法产生一致的估计错误发现的比例。大量的数值实验,沿着一个定量语言学的实证研究,证明了我们的建议相比于许多国家的最先进的方法的优势,当统计中国:新接受的论文(接受作者版本,受英文编辑)
Large-scale multivariate regression is a fundamental statistical tool that finds applications in a wide range of areas. This paper considers the problem of simultaneously testing a large number of general linear hypotheses, encompassing covariate-effect analysis, analysis of variance, and model comparisons. The new challenge that comes along with the overwhelmingly large number of tests is the ubiquitous presence of heavy-tailed and/or highly skewed measurement noise, which is the main reason for the failure of conventional least squares based methods. For large-scale multivariate regression, we develop a set of robust inference methods to explore data features, such as heavy tailedness and skewness, which are invisible to the scope of least squares. The new testing procedure is built on data-adaptive Huber regression, and a new covariance estimator of regression estimates. Under mild conditions, we show that our methods produce consistent estimates of the false discovery proportion. Extensive numerical experiments, along with an empirical study on quantitative linguistics, demonstrate the advantage of our proposal compared to many state-of-the-art methods when the Statistica Sinica: Newly accepted Paper (accepted author-version subject to English editing)