Interactive Q-Learning for Quantiles

Interactive Q-Learning for Quantiles
复制标题

DOI:
10.1080/01621459.2016.1155993
复制
发表时间:
2017-06-01
影响因子:
3.7
通讯作者:
Stefanski, Leonard A.
Stefanski, Leonard A.
中科院分区:
数学1区
文献类型:
--
作者:
Linn, Kristin A.;Laber, Eric B.;Stefanski, Leonard A.

文献摘要

被引文献

相似文献

动态治疗方案是一系列决策规则,每种规则都建议基于患者病史(例如过去的治疗和结果)的特征进行治疗。从数据中估算最佳动态治疗方案的现有方法优化了响应变量的平均值。但是,平均值可能并不总是最合适的绩效摘要。我们得出了针对两阶段二进制治疗环境的响应分布计算的优化概率和分位数的决策规则的估计量。这使得估计动态治疗方案,以优化预先指定点或预先指定的响应分布(例如中位数)的响应累积分布函数。所提出的方法在模拟实验中表现出色。我们用一项依次随机试验的数据说明了我们的方法,其中主要结果是减轻抑郁症状。本文的补充材料可在线获得。
A dynamic treatment regime is a sequence of decision rules, each of which recommends treatment based on features of-patient medical history such as past treatments and outcomes. Existing methods for estimating optimal dynamic treatment regimes from data optimize the mean of a response variable. However, the mean may not always be the most appropriate summary of performance. We derive estimators of decision rules for optimizing probabilities and quantiles computed with respect to the response distribution for two-stage, binary treatment settings. This enables estimation of dynamic treatment regimes that optimize the cumulative distribution function of the response at a prespecified point or a prespecified quantile of the response distribution such as the median. The proposed methods perform favorably in simulation experiments. We illustrate our approach with data from a sequentially randomized trial where the primary outcome is remission of depression symptoms. Supplementary materials for this article are available online.