TARGETED SEQUENTIAL DESIGN FOR TARGETED LEARNING INFERENCE OF THE OPTIMAL TREATMENT RULE AND ITS MEAN REWARD

TARGETED SEQUENTIAL DESIGN FOR TARGETED LEARNING INFERENCE OF THE OPTIMAL TREATMENT RULE AND ITS MEAN REWARD
复制标题

DOI:
10.1214/16-aos1534
复制
发表时间:
2017-12-01
影响因子:
4.5
通讯作者:
van der Laan, Mark J.
van der Laan, Mark J.
中科院分区:
数学1区
文献类型:
--
作者:
Chambaz, Antoine;Zheng, Wenjing;van der Laan, Mark J.

文献摘要

被引文献

相似文献

本文研究了非例外情形下最优治疗规则(TR)及其平均报酬的目标序贯推断问题,即在无基线协变量层的情况下,在伴随边际假设下,我们的关键估计量的定义依赖于目标最小损失估计(TMLE)原则,实际上是在最优TR的当前估计下推断平均奖励。这种数据自适应统计参数本身就值得关注。我们的主要结果是一个中心极限定理,使建设的置信区间的平均回报下,目前估计的最佳TR和最佳TR本身。估计量的渐近方差采用有效影响曲线在极限分布下的方差的形式,从而可以讨论推断的有效性。作为副产品,我们还导出了两个累积伪遗憾的置信区间,这是研究强盗问题的一个关键概念。理论研究的基石之一是关于一致熵积分的鞅的一个新的极大不等式。
This article studies the targeted sequential inference of an optimal treatment rule (TR) and its mean reward in the nonexceptional case, that is, assuming that there is no stratum of the baseline covariates where treatment is neither beneficial nor harmful, and under a companion margin assumption.Our pivotal estimator, whose definition hinges on the targeted minimum loss estimation (TMLE) principle, actually infers the mean reward under the current estimate of the optimal TR. This data-adaptive statistical parameter is worthy of interest on its own. Our main result is a central limit theorem which enables the construction of confidence intervals on both mean rewards under the current estimate of the optimal TR and under the optimal TR itself. The asymptotic variance of the estimator takes the form of the variance of an efficient influence curve at a limiting distribution, allowing to discuss the efficiency of inference.As a by product, we also derive confidence intervals on two cumulated pseudo-regrets, a key notion in the study of bandits problems.A simulation study illustrates the procedure. One of the cornerstones of the theoretical study is a new maximal inequality for martingales with respect to the uniform entropy integral.