Generalization error bounds of dynamic treatment regimes in penalized regression-based learning

Generalization error bounds of dynamic treatment regimes in penalized regression-based learning
复制标题

DOI:
10.1214/22-aos2171
复制
发表时间:
2022-08
期刊:
The Annals of Statistics
影响因子:
--
通讯作者:
E. J. Oh;Min Qian;Y. Cheung
E. J. Oh;Min Qian;Y. Cheung
中科院分区:
其他
文献类型:
--
作者:
E. J. Oh;Min Qian;Y. Cheung

文献摘要

相似文献

动态治疗方案 (DTR) 是一系列决策规则,每个干预阶段都有一个决策规则,它将最新的患者信息映射到推荐的治疗。为给定疾病找到合适的 DTR 是一个具有挑战性的问题,尤其是在观察到大量预后变量时。为了解决这个问题,我们提出了基于惩罚回归的学习方法,具有 l 1 惩罚来估计最佳 DTR,如果实施,该方法将最大化预期结果。我们还提供了在具有多种治疗选项的有限阶段设置中估计 DTR 的泛化误差界限。我们首先检查值和 Q 函数之间的关系,并得出最佳 DTR 和估计 DTR 之间的值差异的有限样本上限。为了实际实现,我们开发了一种通过正交性进行部分正则化的算法来构造最佳 DTR。通过抑郁症临床试验的广泛模拟研究和数据分析证明了所提出方法的优点。
A dynamic treatment regime (DTR) is a sequence of decision rules, one per stage of intervention, that maps up-to-date patient information to a recommended treatment. Discovering an appropriate DTR for a given disease is a challenging issue especially when a large set of prognostic variables are observed. To address this problem, we propose penalized regression-based learning methods with l 1 penalty to estimate the optimal DTR that would maximize the expected outcome if implemented. We also provide generalization error bounds of the estimated DTR in the setting of finite number of stages with multiple treatment options. We first examine the relationship be-tween value and Q-functions and derive a finite sample upper bound on the difference in values between the optimal and the estimated DTRs. For practical implementation, we develop an algorithm with partial regularization via orthogonality to construct the optimal DTR. The advantages of the proposed methods are demonstrated with extensive simulation studies and data analysis of depression clinical trials.