VARIABLE SELECTION FOR REGRESSION MODELS WITH MISSING DATA.

VARIABLE SELECTION FOR REGRESSION MODELS WITH MISSING DATA.
复制标题

DOI:
--
复制
发表时间:
2010
期刊:
影响因子:
1.4
通讯作者:
Ramon I. Garcia;J. Ibrahim;Hong-Tu Zhu
Ramon I. Garcia;J. Ibrahim;Hong-Tu Zhu
中科院分区:
数学3区
文献类型:
--
作者:
Ramon I. Garcia;J. Ibrahim;Hong-Tu Zhu

文献摘要

被引文献

相似文献

本文研究了一类具有缺失数据的统计模型的变量选择问题,其中包括缺失的协变量和/或响应数据。我们研究了平滑剪切绝对偏差惩罚(SCAD)和自适应LASSO,并提出了一个统一的模型选择和估计过程中使用的缺失数据。我们开发了一个计算上有吸引力的算法,同时优化的惩罚似然函数和估计的惩罚参数。特别是,我们建议使用模型选择标准,称为IC(Q)统计量,用于选择惩罚参数。我们表明,基于IC(Q)的变量选择过程自动和一致地选择重要的协变量,并导致有效的估计与预言性质。该方法是非常普遍的,可以适用于许多情况下,涉及缺失的数据,从随机缺失的协变量在任意回归模型中,以nonderably缺失的纵向响应和/或协变量。仿真结果证明了该方法,并研究了有限样本性能的变量选择程序。来自癌症临床试验的黑色素瘤数据被提出来说明所提出的方法。
We consider the variable selection problem for a class of statistical models with missing data, including missing covariate and/or response data. We investigate the smoothly clipped absolute deviation penalty (SCAD) and adaptive LASSO and propose a unified model selection and estimation procedure for use in the presence of missing data. We develop a computationally attractive algorithm for simultaneously optimizing the penalized likelihood function and estimating the penalty parameters. Particularly, we propose to use a model selection criterion, called the IC(Q) statistic, for selecting the penalty parameters. We show that the variable selection procedure based on IC(Q) automatically and consistently selects the important covariates and leads to efficient estimates with oracle properties. The methodology is very general and can be applied to numerous situations involving missing data, from covariates missing at random in arbitrary regression models to nonignorably missing longitudinal responses and/or covariates. Simulations are given to demonstrate the methodology and examine the finite sample performance of the variable selection procedures. Melanoma data from a cancer clinical trial is presented to illustrate the proposed methodology.