A LIKELIHOOD RATIO FRAMEWORK FOR HIGH-DIMENSIONAL SEMIPARAMETRIC REGRESSION

A LIKELIHOOD RATIO FRAMEWORK FOR HIGH-DIMENSIONAL SEMIPARAMETRIC REGRESSION
复制标题

DOI:
10.1214/16-aos1483
复制
发表时间:
2017-12-01
影响因子:
4.5
通讯作者:
Liu, Han
Liu, Han
中科院分区:
数学1区
文献类型:
--
作者:
Ning, Yang;Zhao, Tianqi;Liu, Han

文献摘要

被引文献

相似文献

我们为高维半合理的通用线性模型提供了一个新的推论框架。该框架在高维数据分析中解决了各种具有挑战性的问题,包括不完整的数据,选择偏差和异质性。我们的工作具有三个主要贡献:(i)我们开发了一种正则化统计色谱方法来推断提出的半合理概括性线性模型下感兴趣的参数,而无需估计未知的基础测量函数。 (ii)我们提出了一个基于新的可能性比率的框架来构建高维参数的低维成分的测量后置信区和测试。与现有的调节后推论方法不同,我们的方法基于新颖的定向可能性。 (iii)我们对具有无界内核的U统计量产生了新的浓度不平等和正常近似结果,这些核具有独立感兴趣。我们进一步将理论结果扩展到缺少数据和多个数据集推断的问题。提供了广泛的仿真研究和实际数据分析,以说明所提出的方法。
We propose a new inferential framework for high-dimensional semiparametric generalized linear models. This framework addresses a variety of challenging problems in high-dimensional data analysis, including incomplete data, selection bias and heterogeneity. Our work has three main contributions: (i) We develop a regularized statistical chromatography approach to infer the parameter of interest under the proposed semiparametric generalized linear model without the need of estimating the unknown base measure function. (ii) We propose a new likelihood ratio based framework to construct post-regularization confidence regions and tests for the low dimensional components of high-dimensional parameters. Unlike existing post-regularization inferential methods, our approach is based on a novel directional likelihood. (iii) We develop new concentration inequalities and normal approximation results for U-statistics with unbounded kernels, which are of independent interest. We further extend the theoretical results to the problems of missing data and multiple datasets inference. Extensive simulation studies and real data analysis are provided to illustrate the proposed approach.