A critical appraisal of propensity-score matching in the medical literature between 1996 and 2003

A critical appraisal of propensity-score matching in the medical literature between 1996 and 2003
复制标题

DOI:
10.1002/sim.3150
复制
发表时间:
2008-05-30
影响因子:
2
通讯作者:
Austin, Peter C.
Austin, Peter C.
中科院分区:
医学3区
文献类型:
--
作者:
Austin, Peter C.

文献摘要

被引文献

相似文献

倾向评分方法越来越多地被用于减少治疗选择偏倚在使用观察数据估计治疗效果中的影响。常用的倾向评分方法包括使用倾向评分的协变量调整、倾向评分分层和倾向评分匹配。经验和理论研究表明,倾向评分的匹配比倾向评分的分层消除了更大比例的治疗和未治疗受试者之间的基线差异。然而,倾向分数匹配样本的分析需要适合于匹配对数据的统计方法。我们对1996年至2003年发表在医学文献中的47篇文章进行了批判性评价,这些文章采用了倾向评分匹配。我们发现,只有两篇文章报告了匹配样本中治疗和未治疗受试者之间基线特征的平衡,并使用了正确的统计方法来评估不平衡的程度。13篇文章(28%)在估计治疗效果及其统计学意义时明确使用了适合于匹配数据分析的统计方法。常见错误包括使用对数秩检验比较匹配样本中的Kaplan-Meier生存曲线,在匹配样本中使用考克斯回归、logistic回归、卡方检验、t检验和Wilcoxon秩和检验,从而无法解释数据的匹配性质。我们为采用倾向评分匹配的研究提供了分析和报告指南。版权所有(C)2007约翰威利父子有限公司
Propensity-score methods are increasingly being used to reduce the impact of treatment-selection bias in the estimation of treatment effects using observational data. Commonly used propensity-score methods include covariate adjustment using the propensity score, stratification on the propensity score, and propensity-score matching. Empirical and theoretical research has demonstrated that matching on the propensity score eliminates a greater proportion of baseline differences between treated and untreated subjects than does stratification on the propensity score. However, the analysis of propensity-score-matched samples requires statistical methods appropriate for matched-pairs data. We critically evaluated 47 articles that were published between 1996 and 2003 in the medical literature and that employed propensity-score matching. We found that only two of the articles reported the balance of baseline characteristics between treated and untreated subjects in the matched sample and used correct statistical methods to assess the degree of imbalance. Thirteen (28 per cent) of the articles explicitly used statistical methods appropriate for the analysis of matched data when estimating the treatment effect and its statistical significance. Common errors included using the log-rank test to compare Kaplan-Meier survival curves in the matched sample, using Cox regression, logistic regression, chi-squared tests, t-tests, and Wilcoxon rank sum tests in the matched sample, thereby failing to account for the matched nature of the data. We provide guidelines for the analysis and reporting of studies that employ propensity-score matching. Copyright (C) 2007 John Wiley & Sons, Ltd.