Multiple hypothesis testing in experimental economics

Multiple hypothesis testing in experimental economics
复制标题

DOI:
10.1007/s10683-018-09597-5
复制
发表时间:
2019-12-01
影响因子:
2.3
通讯作者:
Xu, Yang
Xu, Yang
中科院分区:
经济学2区
文献类型:
--
作者:
List, John A.;Shaikh, Azeem M.;Xu, Yang

文献摘要

被引文献

相似文献

经济学实验数据的分析通常涉及同时检验多个零假设。这些不同的零假设在这种情况下自然出现,至少有三个不同的原因:当有多个感兴趣的结果,并希望确定治疗对这些结果中的哪一个有效果时;当治疗的效果可能是异质的,因为它在由观察到的特征定义的亚组之间变化,并希望确定治疗对这些亚组中的哪一个有效果时;最后,当存在多个感兴趣的处理并且希望确定哪些处理相对于对照或相对于每个其它处理具有效果时。在本文中,我们提供了一个基于Bootstrap的程序,同时使用简单的随机抽样的实验数据来测试这些零假设分配治疗状态的单位。使用Romano和Wolf(Ann Stat 38:598-633,2010)中的一般结果,我们在弱假设下证明了我们的方法(1)渐近控制了familywise错误率-一个或多个错误拒绝的概率-并且(2)渐近平衡,因为拒绝任何真实零假设的边际概率在大样本中近似相等。重要的是,通过纳入经典的多重检验程序中忽略的相关信息,如Bonferroni和霍尔姆校正,我们的程序具有更大的能力来检测真正的假零假设。在存在多种治疗的情况下,我们还展示了如何利用零假设的逻辑限制来进一步提高功效。我们通过重新审视Karlan和List(Am Econ Rev 97(5):1774-1793,2007)关于人们为什么向慈善事业捐款的研究来说明我们的方法。
The analysis of data from experiments in economics routinely involves testing multiple null hypotheses simultaneously. These different null hypotheses arise naturally in this setting for at least three different reasons: when there are multiple outcomes of interest and it is desired to determine on which of these outcomes a treatment has an effect; when the effect of a treatment may be heterogeneous in that it varies across subgroups defined by observed characteristics and it is desired to determine for which of these subgroups a treatment has an effect; and finally when there are multiple treatments of interest and it is desired to determine which treatments have an effect relative to either the control or relative to each of the other treatments. In this paper, we provide a bootstrap-based procedure for testing these null hypotheses simultaneously using experimental data in which simple random sampling is used to assign treatment status to units. Using the general results in Romano and Wolf (Ann Stat 38:598-633, 2010), we show under weak assumptions that our procedure (1) asymptotically controls the familywise error rate-the probability of one or more false rejections-and (2) is asymptotically balanced in that the marginal probability of rejecting any true null hypothesis is approximately equal in large samples. Importantly, by incorporating information about dependence ignored in classical multiple testing procedures, such as the Bonferroni and Holm corrections, our procedure has much greater ability to detect truly false null hypotheses. In the presence of multiple treatments, we additionally show how to exploit logical restrictions across null hypotheses to further improve power. We illustrate our methodology by revisiting the study by Karlan and List (Am Econ Rev 97(5):1774-1793, 2007) of why people give to charitable causes.