The likelihood ratio test in high-dimensional logistic regression is asymptotically a rescaled Chi-square

The likelihood ratio test in high-dimensional logistic regression is asymptotically a rescaled Chi-square
复制标题

DOI:
10.1007/s00440-018-00896-9
复制
发表时间:
2017-06
影响因子:
2
通讯作者:
P. Sur;Yuxin Chen;E. Candès
P. Sur;Yuxin Chen;E. Candès
中科院分区:
数学1区
文献类型:
--
作者:
P. Sur;Yuxin Chen;E. Candès

文献摘要

被引文献

相似文献

逻辑回归每天被用于数千次拟合数据,预测未来结果,并评估解释变量的统计意义。当用于统计推断时,logistic模型通过使用似然比检验(LRT)的分布近似值来产生回归系数的ep值。事实上,威尔克斯定理断言,只要我们有一个固定的变量数p,两倍的对数似然比(LLR)在大样本量n的限制下是可变的;这里,是一个卡方,有k个自由度,k是被测试的变量数。在本文中,我们证明了当p与ton相比不可忽略时,Wilks定理不成立,并且卡方近似是非常不正确的;事实上,这种近似产生的p值太小了(在零假设下)。假设n和p以这样的方式变大,对于某个常数。(For,因此LRT在此制度下不感兴趣。)本文证明了对于一类Logistic模型,当维数比为正时,其LLR收敛于一个尺度卡方,即,其中尺度因子大于1。因此,LLR大于经典假设。例如,当。在一般情况下,我们展示了如何计算的比例因子,通过求解两个方程的两个未知数的非线性系统。我们的数学参数涉及和使用的技术,从近似的消息传递理论,从非渐近随机矩阵理论和凸几何。我们还补充我们的数学研究表明,新的极限分布是准确的有限样本量。最后,本文的研究结果也可以推广到其他一些回归模型,如概率单位回归模型。
Logistic regression is used thousands of times a day to fit data, predict future outcomes, and assess the statistical significance of explanatory variables. When used for the purpose of statistical inference, logistic models producep-values for the regression coefficients by using an approximation to the distribution of the likelihood-ratio test (LRT). Indeed, Wilks’ theorem asserts that whenever we have a fixed numberpof variables, twice the log-likelihood ratio (LLR)is distributed as avariable in the limit of large sample sizesn; here,is a Chi-square withkdegrees of freedom andkthe number of variables being tested. In this paper, we prove that whenpis not negligible compared ton, Wilks’ theorem does not hold and that the Chi-square approximation is grossly incorrect; in fact, this approximation producesp-values that are far too small (under the null hypothesis). Assume thatnandpgrow large in such a way thatfor some constant. (For,so that the LRT is not interesting in this regime.) We prove that for a class of logistic models, the LLR converges to arescaledChi-square, namely,, where the scaling factoris greater than one as soon as the dimensionality ratiois positive. Hence, the LLR is larger than classically assumed. For instance, when,. In general, we show how to compute the scaling factor by solving a nonlinear system of two equations with two unknowns. Our mathematical arguments are involved and use techniques from approximate message passing theory, from non-asymptotic random matrix theory and from convex geometry. We also complement our mathematical study by showing that the new limiting distribution is accurate for finite sample sizes. Finally, all the results from this paper extend to some other regression models such as the probit regression model.