Collaborative Double Robust Targeted Maximum Likelihood Estimation

Collaborative Double Robust Targeted Maximum Likelihood Estimation
复制标题

DOI:
10.2202/1557-4679.1181
复制
发表时间:
2010-01-01
影响因子:
1.2
通讯作者:
Gruber, Susan
Gruber, Susan
中科院分区:
数学4区
文献类型:
--
作者:
van der Laan, Mark J.;Gruber, Susan

文献摘要

被引文献

相似文献

协作双稳健目标最大似然估计代表了在半参数模型中数据生成分布的路径可微参数的标准目标最大似然估计的基础上的进一步进步,在货车der Laan,Rubin(2006)中引入。目标最大似然方法涉及波动观测数据密度的相关因子(Q)的初始估计,以便针对感兴趣的参数进行偏差/方差权衡。波动涉及估计可能性g的干扰参数部分。TMLE已被证明是一致的和渐近正态分布(CAN)的规律性条件下,当这两个因素的数据的可能性之一是正确指定,在这篇文章中,我们提供了一个模板,应用合作目标最大似然估计(C-TMLE)的路径可微参数估计半参数,参数模型该程序创建一个序列的候选人有针对性的最大似然估计的基础上的初始估计Q加上一系列越来越多的非参数估计g。在偏离当前技术水平的滋扰参数估计的情况下,g的C-TMLE估计是基于使用滋扰参数来执行波动的相关因子Q的目标最大似然估计器的损失函数而不是滋扰参数本身的损失函数来构造的。基于似然的交叉验证用于在该序列中Q(0)的所有候选TMLE估计量中选择最佳估计量。我们提出了“协作双重稳健性”的理论结果,证明了即使Q和g都被错误指定,协作目标最大似然估计也是CAN的,只要g解决了由Q和真实Q(0)之间的差异所隐含的指定得分方程。本文还建立了目标参数的C-DR-TMLE的渐近线性定理,证明了C-DR-TMLE对真值的适应性更强,并由此证明了C-DR-TMLE对真值的适应性更强。如果第一阶段密度估计器本身相对于目标参数做得很好,那么它甚至可以是超高效的。这项研究提供了一个模板,用于对大型(无限维)半参数模型中数据概率分布的特定目标特征进行有针对性的高效和鲁棒的基于损失的学习,同时仍提供置信区间和p值方面的统计推断。这项研究也打破了一个禁忌(例如,在因果推断领域的倾向评分文献中)关于使用似然的相关部分来微调滋扰参数/删失机制/处理机制的拟合。
Collaborative double robust targeted maximum likelihood estimators represent a fundamental further advance over standard targeted maximum likelihood estimators of a pathwise differentiable parameter of a data generating distribution in a semiparametric model, introduced in van der Laan, Rubin (2006). The targeted maximum likelihood approach involves fluctuating an initial estimate of a relevant factor (Q) of the density of the observed data, in order to make a bias/variance tradeoff targeted towards the parameter of interest. The fluctuation involves estimation of a nuisance parameter portion of the likelihood, g. TMLE has been shown to be consistent and asymptotically normally distributed (CAN) under regularity conditions, when either one of these two factors of the likelihood of the data is correctly specified, and it is semiparametric efficient if both are correctly specified.In this article we provide a template for applying collaborative targeted maximum likelihood estimation (C-TMLE) to the estimation of pathwise differentiable parameters in semi-parametric models. The procedure creates a sequence of candidate targeted maximum likelihood estimators based on an initial estimate for Q coupled with a succession of increasingly non-parametric estimates for g. In a departure from current state of the art nuisance parameter estimation, C-TMLE estimates of g are constructed based on a loss function for the targeted maximum likelihood estimator of the relevant factor Q that uses the nuisance parameter to carry out the fluctuation, instead of a loss function for the nuisance parameter itself. Likelihood-based cross-validation is used to select the best estimator among all candidate TMLE estimators of Q(0) in this sequence. A penalized-likelihood loss function for Q is suggested when the parameter of interest is borderline-identifiable.We present theoretical results for "collaborative double robustness," demonstrating that the collaborative targeted maximum likelihood estimator is CAN even when Q and g are both misspecified, providing that g solves a specified score equation implied by the difference between the Q and the true Q(0). This marks an improvement over the current definition of double robustness in the estimating equation literature.We also establish an asymptotic linearity theorem for the C-DR-TMLE of the target parameter, showing that the C-DR-TMLE is more adaptive to the truth, and, as a consequence, can even be super efficient if the first stage density estimator does an excellent job itself with respect to the target parameter.This research provides a template for targeted efficient and robust loss-based learning of a particular target feature of the probability distribution of the data within large (infinite dimensional) semi-parametric models, while still providing statistical inference in terms of confidence intervals and p-values. This research also breaks with a taboo (e.g., in the propensity score literature in the field of causal inference) on using the relevant part of likelihood to fine-tune the fitting of the nuisance parameter/censoring mechanism/treatment mechanism.