Ambiguous Dynamic Treatment Regimes: A Reinforcement Learning Approach

Ambiguous Dynamic Treatment Regimes: A Reinforcement Learning Approach
复制标题

模糊动态治疗方案:强化学习方法

DOI:
10.2139/ssrn.3980837
复制
发表时间:
2021
期刊:
ArXiv
影响因子:
--
通讯作者:
S. Saghafian
S. Saghafian
中科院分区:
--
文献类型:
--
作者:
S. Saghafian

文献摘要

参考文献

被引文献

相似文献

各种研究的一个主要研究目标是使用观察数据集,并提供一套新的反事实指导方针,可以产生因果改进。动态治疗方案(DTRs)被广泛研究,以使这一过程形式化,并使研究人员能够找到既个性化又动态的指导方针。然而,寻找最佳dtr的现有方法往往依赖于在现实应用(例如,医疗决策或公共政策)中违反的假设,特别是当(a)存在不可忽视的未观察到的混杂因素,以及(b)未观察到的混杂因素是时变的(例如,受先前行为的影响)。当这些假设被违反时,人们经常面临关于获得最优DTR所需假设的潜在因果模型的模糊性。这种模糊性是不可避免的,因为无法从观察到的数据中理解未观察到的混杂因素的动态及其对观察到的数据部分的因果影响。在我们的合作医院(梅奥诊所)为移植术后面临新发糖尿病的患者寻找更好的治疗方案的案例研究的激励下,我们将dtr扩展到一个新的类别,称为模糊动态治疗方案(adtr),其中治疗方案的因果影响是基于潜在因果模型的“云”来评估的。然后,我们将adtr与Saghafian(2018)提出的模糊部分可观察马尔可夫决策过程(apomdp)联系起来,并将未观察到的混杂因素视为潜在变量,但对观察到的变量具有模糊的动态和因果关系。利用这种联系,我们开发了两种强化学习方法,称为直接增强v -学习(davi - learning)和安全增强v -学习(SAV-Learning),它们能够使用观察到的数据有效地学习最佳治疗方案。我们建立了这些学习方法的理论结果,包括(弱)一致性和渐近正态性。我们在案例研究(使用临床数据)和模拟实验(使用合成数据)中进一步评估了这些学习方法的性能。我们发现了我们提出的方法的有希望的结果,表明它们甚至与一个既知道真实因果模型(数据生成过程)又知道该模型下的最佳制度的假想预言者相比表现良好。最后,我们强调,我们的方法可以实现双向个性化;获得的治疗方案可以根据患者的特点和医生的偏好进行个性化。这篇论文被数据科学的David Simchi-Levi接受。补充材料:数据文件和在线附录可在https://doi.org/10.1287/mnsc.2022.00883上获得。
A main research goal in various studies is to use an observational data set and provide a new set of counterfactual guidelines that can yield causal improvements. Dynamic Treatment Regimes (DTRs) are widely studied to formalize this process and enable researchers to find guidelines that are both personalized and dynamic. However, available methods in finding optimal DTRs often rely on assumptions that are violated in real-world applications (e.g., medical decision making or public policy), especially when (a) the existence of unobserved confounders cannot be ignored, and (b) the unobserved confounders are time varying (e.g., affected by previous actions). When such assumptions are violated, one often faces ambiguity regarding the underlying causal model that is needed to be assumed to obtain an optimal DTR. This ambiguity is inevitable because the dynamics of unobserved confounders and their causal impact on the observed part of the data cannot be understood from the observed data. Motivated by a case study of finding superior treatment regimes for patients who underwent transplantation in our partner hospital (Mayo Clinic) and faced a medical condition known as new-onset diabetes after transplantation, we extend DTRs to a new class termed Ambiguous Dynamic Treatment Regimes (ADTRs), in which the causal impact of treatment regimes is evaluated based on a “cloud” of potential causal models. We then connect ADTRs to Ambiguous Partially Observable Markov Decision Processes (APOMDPs) proposed by Saghafian (2018) , and consider unobserved confounders as latent variables but with ambiguous dynamics and causal effects on observed variables. Using this connection, we develop two reinforcement learning methods termed Direct Augmented V-Learning (DAV-Learning) and Safe Augmented V-Learning (SAV-Learning), which enable using the observed data to effectively learn an optimal treatment regime. We establish theoretical results for these learning methods, including (weak) consistency and asymptotic normality. We further evaluate the performance of these learning methods both in our case study (using clinical data) and in simulation experiments (using synthetic data). We find promising results for our proposed approaches, showing that they perform well even compared with an imaginary oracle who knows both the true causal model (of the data-generating process) and the optimal regime under that model. Finally, we highlight that our approach enables a two-way personalization; obtained treatment regimes can be personalized based on both patients’ characteristics and physicians’ preferences. This paper was accepted by David Simchi-Levi, data science. Supplemental Material: The data files and online appendix are available at https://doi.org/10.1287/mnsc.2022.00883 .
创新的医疗保健服务:设计移动医疗干预措施的科学和监管挑战。
DOI: 10.31478/202108b
发表时间: 2021
期刊: NAM perspectives
影响因子: --
作者:
Saghafian,Soroush;Murphy,SusanA
通讯作者: Murphy,SusanA
无限视野强化学习中的混杂鲁棒策略评估
DOI: --
发表时间: 2020
期刊: Advances in neural information processing systems
影响因子: --
作者:
Kallus, Nathan;Zhou, Angela
通讯作者: Zhou, Angela
近端强化学习:部分观察马尔可夫决策过程中的高效离策略评估
DOI: 10.1287/opre.2021.0781
发表时间: 2023
影响因子: 2.7
作者:
Bennett, Andrew;Kallus, Nathan
通讯作者: Kallus, Nathan
DOI: 10.1080/01621459.2017.1330204
发表时间: 2018
影响因子: 3.7
作者:
Wang L;Zhou Y;Song R;Sherwood B
通讯作者: Sherwood B
DOI: 10.1080/01621459.2020.1831925
发表时间: 2020-11-28
影响因子: 3.7
作者:
Nie, Xinkun;Brunskill, Emma;Wager, Stefan
通讯作者: Wager, Stefan