Dimension reduction and variable selection in case control studies via regularized likelihood optimization

Dimension reduction and variable selection in case control studies via regularized likelihood optimization
复制标题

DOI:
10.1214/09-ejs537
复制
发表时间:
2009-01-01
影响因子:
1.1
通讯作者:
Barbu, Adrian
Barbu, Adrian
中科院分区:
数学3区
文献类型:
--
作者:
Buneo, Florentino;Barbu, Adrian

文献摘要

被引文献

相似文献

降维和变量选择在病例对照研究中是常规进行的,但是关于所得到的估计的理论方面的文献很少。我们把我们的贡献,这方面的文献,通过研究估计得到l(1)惩罚似然优化。我们证明了l(1)惩罚的回顾似然的最优解与l(1)惩罚的前瞻似然的最优解是一致的。这扩展了普伦蒂斯和派克(1979)的结果,得到了非正则化的似然。我们建立了模型选择后优势比的超范数一致性和估计量子集选择的一致性。我们的理论结果的新奇在于这些属性的病例对照抽样方案下的研究。我们的结果举行的选择在一个大的候选变量集合,允许基数依赖,并大于样本量。我们补充我们的理论结果与一种新的方法来确定数据驱动的调谐参数,基于二分法。由此产生的程序提供了显着的计算节省相比,基于网格搜索的方法。所有的数值实验都有力地支持了我们的理论发现。
Dimension reduction and variable selection are performed routinely in case-control studies, but the literature on the theoretical aspects of the resulting estimates is scarce. We bring our contribution to this literature by studying estimators obtained via l(1) penalized likelihood optimization. We show that the optimizers of the l(1) penalized retrospective likelihood coincide with the optimizers of the l(1) penalized prospective likelihood. This extends the results of Prentice and Pyke (1979), obtained for non-regularized likelihoods. We establish both the sup-norm consistency of the odds ratio, after model selection, and the consistency of subset selection of our estimators. The novelty of our theoretical results consists in the study of these properties under the case-control sampling scheme. Our results hold for selection performed over a large collection of candidate variables, with cardinality allowed to depend and be greater than the sample size. We complement our theoretical results with a novel approach of determining data driven tuning parameters, based on the bisection method. The resulting procedure offers significant computational savings when compared with grid search based methods. All our numerical experiments support strongly our theoretical findings.