Propensity score and proximity matching using random forest.

Propensity score and proximity matching using random forest.
复制标题

DOI:
10.1016/j.cct.2015.12.012
复制
发表时间:
2016-03
影响因子:
2.2
通讯作者:
Fan J
Fan J
中科院分区:
医学4区
文献类型:
--
作者:
Zhao P;Su X;Ge T;Fan J

文献摘要

相似文献

为了从观察数据中得出无偏推断,通常采用匹配方法来产生所有背景变量方面的平衡治疗组和对照组。倾向得分一直是这一研究领域的关键组成部分。然而,文献中基于倾向评分的匹配方法有几个局限性,如模型误设定,分类变量超过两个水平,处理缺失数据的困难,以及非线性关系。随机森林,平均结果从许多决策树,是非参数的性质,直接使用,并能够解决这些问题。更重要的是,随机森林提供的精确度可以为我们提供更准确和更少依赖模型的倾向评分估计。此外,邻近矩阵,随机森林的副产品,可以自然地用作可以在匹配中使用的观测之间的距离度量。建议随机森林匹配方法适用于全国健康和营养调查(NHANES)的数据。我们的研究结果表明,所提出的方法可以产生很好的平衡的治疗和控制组。最后通过实例说明了该方法可以有效地处理协变量中的缺失数据。
In order to derive unbiased inference from observational data, matching methods are often applied to produce balanced treatment and control groups in terms of all background variables. Propensity score has been a key component in this research area. However, propensity score based matching methods in the literature have several limitations, such as model misspecifications, categorical variables with more than two levels, difficulties in handling missing data, and nonlinear relationships. Random forest, averaging outcomes from many decision trees, is nonparametric in nature, straightforward to use, and capable of solving these issues. More importantly, the precision afforded by random forest may provide us with a more accurate and less model dependent estimate of the propensity score. In addition, the proximity matrix, a by-product of the random forest, may naturally serve as a distance measure between observations that can be used in matching. The proposed random forest based matching methods are applied to data from the National Health and Nutrition Examination Survey (NHANES). Our results show that the proposed methods can produce well balanced treatment and control groups. An illustration is also provided that the methods can effectively deal with missing data in covariates.