Monte Carlo EM for missing covariates in parametric regression models

Monte Carlo EM for missing covariates in parametric regression models
复制标题

DOI:
10.1111/j.0006-341x.1999.00591.x
复制
发表时间:
1999-06-01
期刊:
影响因子:
1.9
通讯作者:
Lipsitz, SR
Lipsitz, SR
中科院分区:
数学3区
文献类型:
--
作者:
Ibrahim, JG;Chen, MH;Lipsitz, SR

文献摘要

被引文献

相似文献

本文提出了一种估计具有任意数目缺失协变量的一般参数回归模型的参数的方法。我们允许任何缺失数据模式,并假设缺失数据机制自始至终都是可以忽略的。当缺失的协变量是分类的时,用于获得参数估计值的有用技术是通过Ibrahim(1990,Journal of the American Statistical Association 85,765-769)中提出的权重方法的EM算法。我们将这种方法扩展到连续或混合分类和连续协变量,并为任意参数回归模型,通过调整的EM算法的Monte Carlo版本,如Wei和坦纳(1990年,美国,统计协会85,699-704)。此外,我们讨论了吉布斯抽样的条件分布的缺失的协变量给定的观测数据的抽样,并表明适当的完整条件是对数凹的。条件分布的对数-对数特性将有助于通过Gilks和Wild(1992,Applied Statistics 41,337-348)的自适应拒绝算法直接实现吉布斯采样器。我们假设给定协变量的响应模型是任意参数回归模型,例如广义线性模型、参数生存模型或非线性模型。我们将协变量的边际分布建模为一维条件分布的乘积。这使我们在对协变量分布进行建模时具有很大的灵活性,并减少了E步骤中引入的多余参数的数量。我们提出的例子涉及模拟和真实的数据。
We propose a method for estimating parameters for general parametric regression models with an arbitrary number of missing covariates. We allow any pattern of missing data and assume that the missing data mechanism is ignorable throughout. When the missing covariates are categorical, a useful technique for obtaining parameter estimates is the EM algorithm by the method of weights proposed in Ibrahim (1990, Journal of the American Statistical Association 85, 765-769). We extend this method to continuous or mixed categorical and continuous covariates, and for arbitrary parametric regression models, by adapting a Monte Carlo version of the EM algorithm as discussed by Wei and Tanner (1990, Journal of the American, Statistical Association 85, 699-704). In addition, we discuss the Gibbs sampler for sampling from the conditional distribution of the missing covariates given the observed data and show that the appropriate complete conditionals are log-concave. The log-concavity property of the conditional distributions will facilitate a straightforward implementation of the Gibbs sampler via the adaptive rejection algorithm of Gilks and Wild (1992, Applied Statistics 41, 337-348). We assume the model for the response given the covariates is an arbitrary parametric regression model, such as a generalized linear model, a parametric survival model, or a nonlinear model. We model the marginal distribution of the covariates as a product of one-dimensional conditional distributions. This allows us a great deal of flexibility in modeling the distribution of the covariates and reduces the number of nuisance parameters that are introduced in the E-step. We present examples involving both simulated and real data.