New Statistical Learning Methods for Estimating Optimal Dynamic Treatment Regimes.

New Statistical Learning Methods for Estimating Optimal Dynamic Treatment Regimes.
复制标题

DOI:
10.1080/01621459.2014.937488
复制
发表时间:
2015
影响因子:
3.7
通讯作者:
Kosorok MR
Kosorok MR
中科院分区:
数学1区
文献类型:
--
作者:
Zhao YQ;Zeng D;Laber EB;Kosorok MR

文献摘要

被引文献

相似文献

动态治疗方案(DTR)是针对个体患者的连续决策规则,可以随着时间的推移适应不断演变的疾病。目标是适应患者之间的异质性,并找到如果实施DTR将产生最佳长期结果的DTR。我们介绍了两种新的估计最优DTR的统计学习方法,称为后向结果加权学习(Bowl)和同时结果加权学习(SOWL)。这些方法将个性化的治疗选择转化为顺序或同时分类问题,因此可以通过修改现有的机器学习技术来应用。所提出的方法是基于在所有DRR上直接最大化预期长期结果的非参数估计;这与基于回归的方法(例如Q-学习)本质上不同,后者间接地尝试这种最大化并严重依赖假设回归模型的正确性。我们证明了所得规则是一致的,并利用估计规则提供了误差的有限样本界。仿真结果表明,与Q学习方法相比,所提出的方法具有更好的动态反馈性能,尤其是在小样本情况下。我们使用一项戒烟临床试验的数据来说明这些方法。
Dynamic treatment regimes (DTRs) are sequential decision rules for individual patients that can adapt over time to an evolving illness. The goal is to accommodate heterogeneity among patients and find the DTR which will produce the best long term outcome if implemented. We introduce two new statistical learning methods for estimating the optimal DTR, termed backward outcome weighted learning (BOWL), and simultaneous outcome weighted learning (SOWL). These approaches convert individualized treatment selection into an either sequential or simultaneous classification problem, and can thus be applied by modifying existing machine learning techniques. The proposed methods are based on directly maximizing over all DTRs a nonparametric estimator of the expected long-term outcome; this is fundamentally different than regression-based methods, for example Q-learning, which indirectly attempt such maximization and rely heavily on the correctness of postulated regression models. We prove that the resulting rules are consistent, and provide finite sample bounds for the errors using the estimated rules. Simulation results suggest the proposed methods produce superior DTRs compared with Q-learning especially in small samples. We illustrate the methods using data from a clinical trial for smoking cessation.