Can Nonrandomized Experiments Yield Accurate Answers? A Randomized Experiment Comparing Random and Nonrandom Assignments

Can Nonrandomized Experiments Yield Accurate Answers? A Randomized Experiment Comparing Random and Nonrandom Assignments
复制标题

DOI:
10.1198/016214508000000733
复制
发表时间:
2008-12-01
影响因子:
3.7
通讯作者:
Steiner, Peter M.
Steiner, Peter M.
中科院分区:
数学1区
文献类型:
--
作者:
Shadish, William R.;Clark, M. H.;Steiner, Peter M.

文献摘要

被引文献

相似文献

使用非随机实验的一个关键理由是,经过适当的调整,他们的结果可以很好地接近随机实验的结果。这一假说并没有得到实证研究的一致支持;然而,以前用于研究这一假说的方法混淆了分配方法和其他研究特征。为了避免这些混杂因素,本研究随机分配参与者参加随机实验或非随机实验。在随机实验中,参与者被随机分配到数学或词汇训练中;在非随机实验中,参与者选择他们的训练。这项研究保持了实验的所有其他特征不变:它仔细地测量了可能预测参与者选择的条件的预测变量,并对所有参与者的词汇和数学结果进行了测量。以协变量调整后的随机化结果为基准,普通线性回归将非随机化实验中的偏差降低了84-94%。倾向分数分层、加权和协方差调整减少了约58%-96%的偏差。这取决于结果的衡量和调整方法。当根据便利性的预测因素(性别、年龄、婚姻状况和种族)而不是根据可能包括这些因素的更广泛的预测因素构建得分时,倾向性得分调整的效果很差。
A key justification for using nonrandomized experiments is that, with proper adjustment, their results can well approximate results from randomized experiments. This hypothesis has not been consistently supported by empirical studies; however, previous methods used to study this hypothesis have confounded assignment method with other study features. To avoid these confounding factors, this study randomly assigned participants to be in a randomized experiment or a nonrandomized experiment. In the randomized experiment, participants were randomly assigned to mathematics or vocabulary training; in the nonrandomized experiment, participants chose their training. The study held all other features of the experiment constant: it carefully measured pretest variables that might predict the condition that participants chose, and all participants were measured on vocabulary and mathematics outcomes. Ordinary linear regression reduced bias in the nonrandomized experiment by 84-94% using covariate-adjusted randomized results as the benchmark. Propensity score stratification, weighting, and covariance adjustment reduced bias by about 58-96%. depending on the outcome measure and adjustment method. Propensity score adjustment performed poorly when the scores were constructed from predictors of convenience (sex, age, marital status, and ethnicity) rather than from a broader set of predictors that might include these.