Likelihood Methods for Regression Models with Expensive Variables Missing by Design

Likelihood Methods for Regression Models with Expensive Variables Missing by Design
复制标题

DOI:
10.1002/bimj.200810487
复制
发表时间:
2009-02-01
影响因子:
1.7
通讯作者:
McLeish, Donald L.
McLeish, Donald L.
中科院分区:
生物学3区
文献类型:
--
作者:
Zhao, Yang;Lawless, Jerald F.;McLeish, Donald L.

文献摘要

被引文献

相似文献

在一些涉及回归的应用中,某些变量的值因设计而为某些个体丢失。例如,在两阶段的研究中(Zhao和Lipsitz, 1992),在第一阶段的随机样本中收集“便宜”变量的数据,然后在第二阶段的子样本中测量“昂贵”变量。因此,“昂贵的”变量在第一阶段的设计中丢失了。对于协变量或响应丢失的情况,已经提出了估计函数和似然方法。我们扩展了半参数最大似然(SPML)方法用于缺失协变量问题(例如Chen, 2004; Ibrahim等人,2005;Zhang和Rockette, 2005, 2007),以处理更一般的情况,其中协变量和/或响应因设计而缺失,并表明轮廓似然比检验和区间估计很容易实现。提供了模拟研究来检查似然方法的性能,并将其与涉及(a)缺少协变量和(b)缺少响应变量的问题的估计函数方法的效率进行比较。我们说明了SPML的易于实现,并证明了它的高效率。
In some applications involving regression the values of certain variables are missing by design for some individuals. For example, in two-stage studies (Zhao and Lipsitz, 1992), data on "cheaper" variables are collected on a random sample of individuals in stage 1, and then "expensive" variables are measured for a subsample of these in stage II. So the "expensive" variables are missing by design at stage I. Both estimating function and likelihood methods have been proposed for cases where either covariates or responses are missing. We extend the semiparametric maximum likelihood (SPML) method for missing covariate problems (e.g. Chen, 2004; Ibrahim et al., 2005; Zhang and Rockette, 2005, 2007) to deal with more general cases where covariates and/or responses are missing by design, and show that profile likelihood ratio tests and interval estimation are easily implemented. Simulation studies are provided to examine the performance of the likelihood methods and to compare their efficiencies with estimating function methods for problems involving (a) a missing covariate and (b) a missing response variable. We illustrate the ease of implementation of SPML and demonstrate its high efficiency.