On Nearly Assumption-Free Tests of Nominal Confidence Interval Coverage for Causal Parameters Estimated by Machine Learning

On Nearly Assumption-Free Tests of Nominal Confidence Interval Coverage for Causal Parameters Estimated by Machine Learning
复制标题

DOI:
10.1214/20-sts786
复制
发表时间:
2020-08-01
影响因子:
5.7
通讯作者:
Robins, James M.
Robins, James M.
中科院分区:
数学2区
文献类型:
--
作者:
Liu, Lin;Mukherjee, Rajarshi;Robins, James M.

文献摘要

被引文献

相似文献

对于许多感兴趣的因果效应参数,双重鲁棒机器学习(DRML)估计(psi)在cap(1)上是最先进的,结合了机器学习的良好预测性能;双重鲁棒估计的偏差降低;以及交叉拟合样本分裂的分析易处理性和偏差降低。尽管如此,即使没有不可测量因素的干扰,(1 - alpha)Wald置信区间(psi)超过上限(1)+/- z(alpha/2)盖上的(s.e)[盖(1)上的(psi)]即使在大样品中也可能仍然是下盖,因为(psi)在盖(1)上的偏差可能与其阶数n(-1/2)的标准误差具有相同或甚至更大的阶数。在本文中,我们引入了本质上无干扰的检验,其(i)可以证伪零假设,即(psi)对cap(1)的偏差具有比其标准误差更小的阶,(ii)可以提供Wald区间的真实覆盖率的置信上限,以及(iii)在对干扰参数没有平滑性/稀疏性假设的情况下在零值下有效。我们称之为无假设经验覆盖检验(AFECTs)的检验基于U统计量,该统计量估计了(psi)相对于cap(1)的部分偏差。首先,包括我们在内的零假设检验(即偏差与其标准误差之比小于某个阈值δ)不可能是一致的[没有额外的假设(例如,平滑度或稀疏度)可能不正确]。其次,上述权利要求仅适用于特定类别中的某些参数。对于大多数其他人来说,我们的结果显然不那么尖锐。特别地,对于这些参数,我们不能直接测试标称Wald间隔(psi)over cap(1)+/- z(alpha/2)(s.e)over cap[(psi)over cap(1)]是否覆盖。然而,我们经常可以测试分析师使用的平滑性和/或稀疏性假设的有效性,以证明报告的Wald区间的实际覆盖率不小于标称值。第三,在正文中,除了第1节中的模拟研究之外,我们假设我们处于半监督数据设置中(其中存在一个更大的数据集,仅包含协变量的信息),允许我们将协变量的协方差矩阵视为已知。在第1节的模拟中,我们考虑需要估计协方差矩阵的设置。在模拟中,我们使用了一个数据自适应估计器,它在我们的模拟中表现得很好,但估计器的理论采样行为仍然未知。
For many causal effect parameters of interest, doubly robust machine learning (DRML) estimators (psi) over cap (1) are the state-of-the-art, incorporating the good prediction performance of machine learning; the decreased bias of doubly robust estimators; and the analytic tractability and bias reduction of sample splitting with cross-fitting. Nonetheless, even in the absence of confounding by unmeasured factors, the nominal (1 - alpha) Wald confidence interval (psi) over cap (1) +/- z(alpha/2)(s.e) over cap[(psi) over cap (1)] may still undercover even in large samples, because the bias of (psi) over cap (1) may be of the same or even larger order than its standard error of order n(-1/2).In this paper, we introduce essentially assumption-free tests that (i) can falsify the null hypothesis that the bias of (psi) over cap (1) is of smaller order than its standard error, (ii) can provide a upper confidence bound on the true cover-age of the Wald interval, and (iii) are valid under the null under no smoothness/sparsity assumptions on the nuisance parameters. The tests, which we refer to as Assumption Free Empirical Coverage Tests (AFECTs), are based on a U-statistic that estimates part of the bias of (psi) over cap (1).Our claims need to be tempered in several important ways. First no test, including ours, of the null hypothesis that the ratio of the bias to its standard error is smaller than some threshold delta can be consistent [without additional assumptions (e.g., smoothness or sparsity) that may be incorrect]. Second, the above claims only apply to certain parameters in a particular class. For most of the others, our results are unavoidably less sharp. In particular, for these parameters, we cannot directly test whether the nominal Wald interval (psi) over cap (1) +/- z(alpha/2)(s.e) over cap[(psi) over cap (1)] undercovers. However, we can often test the validity of the smoothness and/or sparsity assumptions used by an analyst to justify a claim that the reported Wald interval's actual coverage is no less than nominal. Third, in the main text, with the exception of the simulation study in Sec- tion 1, we assume we are in the semisupervised data setting (wherein there is a much larger dataset with information only on the covariates), allowing us to regard the covariance matrix of the covariates as known. In the simulation in Section 1, we consider the setting in which estimation of the covariance ma- trix is required. In the simulation, we used a data adaptive estimator which performs very well in our simulations, but the estimator's theoretical sampling behavior remains unknown.