Did we personalize? Assessing personalization by an online reinforcement learning algorithm using resampling

Did we personalize? Assessing personalization by an online reinforcement learning algorithm using resampling
复制标题

我们个性化了吗?

DOI:
10.48550/arxiv.2304.05365
复制
发表时间:
2023
期刊:
Mach. Learn.
影响因子:
--
通讯作者:
Susan A. Murphy
Susan A. Murphy
中科院分区:
--
文献类型:
--
作者:
Susobhan Ghosh;Raphael Kim;Prasidh Chhabria;Raaz Dwivedi;Predrag Klasjna;Peng Liao;Kelly W. Zhang;Susan A. Murphy

文献摘要

参考文献

被引文献

相似文献

人们越来越有兴趣使用强化学习(RL)来个性化数字健康中的治疗序列,以支持用户采取更健康的行为。这种顺序决策问题涉及基于用户的上下文(例如,先前的活动水平、位置等)。在线RL是解决这个问题的一种很有前途的数据驱动方法,因为它基于每个用户的历史响应进行学习,并使用这些知识来个性化这些决策。然而,为了决定RL算法是否应该被包括在现实世界部署的“优化”干预中,我们必须评估数据证据,表明RL算法实际上是为其用户提供个性化治疗。由于RL算法中的随机性,人们可能会得到一个错误的印象,即它正在某些状态下学习,并使用这种学习来提供特定的处理。我们使用的个性化的工作定义,并介绍了一个基于重采样的方法,调查是否表现出的RL算法的个性化是一个工件的RL算法的随机性。我们通过分析一项名为HeartSteps的体力活动临床试验的数据来说明我们的方法,该试验包括使用在线RL算法。我们展示了我们的方法如何增强数据驱动的算法个性化广告的真实性,无论是在所有用户中,以及在研究中的特定用户。
There is a growing interest in using reinforcement learning (RL) to personalize sequences of treatments in digital health to support users in adopting healthier behaviors. Such sequential decision-making problems involve decisions about when to treat and how to treat based on the user's context (e.g., prior activity level, location, etc.). Online RL is a promising data-driven approach for this problem as it learns based on each user's historical responses and uses that knowledge to personalize these decisions. However, to decide whether the RL algorithm should be included in an ``optimized'' intervention for real-world deployment, we must assess the data evidence indicating that the RL algorithm is actually personalizing the treatments to its users. Due to the stochasticity in the RL algorithm, one may get a false impression that it is learning in certain states and using this learning to provide specific treatments. We use a working definition of personalization and introduce a resampling-based methodology for investigating whether the personalization exhibited by the RL algorithm is an artifact of the RL algorithm stochasticity. We illustrate our methodology with a case study by analyzing the data from a physical activity clinical trial called HeartSteps, which included the use of an online RL algorithm. We demonstrate how our approach enhances data-driven truth-in-advertising of algorithm personalization both across all users as well as within specific users in the study.
DOI: --
发表时间: 2021-06
期刊: Advances in neural information processing systems
影响因子: --
作者:
Aurélien F. Bibaut;A. Chambaz;Maria Dimakopoulou;Nathan Kallus;M. Laan
通讯作者: Aurélien F. Bibaut;A. Chambaz;Maria Dimakopoulou;Nathan Kallus;M. Laan
DOI: --
发表时间: 2020-02
期刊: Advances in neural information processing systems
影响因子: --
作者:
Kelly W. Zhang;Lucas Janson;S. Murphy
通讯作者: Kelly W. Zhang;Lucas Janson;S. Murphy
DOI: 10.1016/j.cct.2022.107029
发表时间: 2023
影响因子: 2.2
作者:
Forman,EvanM;Berry,MichaelP;Butryn,MeghanL;Hagerman,CharlotteJ;Huang,Zhuoran;Juarascio,AdrienneS;LaFata,EricaM;Ontañón,Santiago;Tilford,JMick;Zhang,Fengqing
通讯作者: Zhang,Fengqing
DOI: 10.1080/01621459.2022.2096620
发表时间: 2021-08
影响因子: 3.7
作者:
Pratik Ramprasad;Yuantong Li;Zhuoran Yang;Zhaoran Wang;W. Sun;Guang Cheng
通讯作者: Pratik Ramprasad;Yuantong Li;Zhuoran Yang;Zhaoran Wang;W. Sun;Guang Cheng
DOI: 10.1093/biomet/asaa070
发表时间: 2021-09
期刊: Biometrika
影响因子: 2.7
作者:
Qian T;Yoo H;Klasnja P;Almirall D;Murphy SA
通讯作者: Murphy SA