Penalized Q-Learning for Dynamic Treatment Regimens.

Penalized Q-Learning for Dynamic Treatment Regimens.
复制标题

DOI:
10.5705/ss.2012.364
复制
发表时间:
2015-07
期刊:
影响因子:
1.4
通讯作者:
Kosorok MR
Kosorok MR
中科院分区:
数学3区
文献类型:
--
作者:
Song R;Wang W;Zeng D;Kosorok MR

文献摘要

被引文献

相似文献

动态治疗方案结合了专门设计的临床试验的累积信息和长期治疗效果。随着这些试验与临床研究的纵向数据结合变得越来越流行,最佳动态治疗方案的统计推断的发展是一个高度优先事项。在本文中,我们提出了一个新的机器学习框架,称为惩罚Q学习,在此框架下建立了有效的统计推断。我们还提出了一个新的统计过程:个人选择和相应的方法,将个人选择惩罚Q-学习。大量的数值研究,提出了与现有的方法相比,在各种情况下,所提出的方法,并证明了所提出的方法是既inquiry和计算上级。这是说明与抑郁症的临床试验研究。
A dynamic treatment regimen incorporates both accrued information and long-term effects of treatment from specially designed clinical trials. As these trials become more and more popular in conjunction with longitudinal data from clinical studies, the development of statistical inference for optimal dynamic treatment regimens is a high priority. In this paper, we propose a new machine learning framework called penalized Q-learning, under which valid statistical inference is established. We also propose a new statistical procedure: individual selection and corresponding methods for incorporating individual selection within penalized Q-learning. Extensive numerical studies are presented which compare the proposed methods with existing methods, under a variety of scenarios, and demonstrate that the proposed approach is both inferentially and computationally superior. It is illustrated with a depression clinical trial study.