HOPE: Human-Centric Off-Policy Evaluation for E-Learning and Healthcare

HOPE: Human-Centric Off-Policy Evaluation for E-Learning and Healthcare
复制标题

DOI:
10.48550/arxiv.2302.09212
复制
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Ge Gao;Song Ju;Markel Sanz Ausin;Min Chi
Ge Gao;Song Ju;Markel Sanz Ausin;Min Chi
中科院分区:
其他
文献类型:
--
作者:
Ge Gao;Song Ju;Markel Sanz Ausin;Min Chi

文献摘要

相似文献

强化学习(RL)已被广泛研究,用于增强各种以人为中心的任务中的人与环境交互,包括电子学习和医疗保健。由于在线部署和评估政策在这些任务中具有很高的风险,因此政策外评估(OPE)对于诱导有效的政策至关重要。然而,在以人为中心的环境中,OPE是具有挑战性的,因为底层状态通常是不可观察的,而只能观察到总体奖励(学生的考试成绩或患者最终是否出院)。在这项工作中,我们提出了一个以人为中心的OPE(希望),以处理部分可观察性和聚合奖励在这样的环境中。具体来说,我们重建即时回报的总回报,考虑部分可观察性估计预期的总回报。我们为所提出的方法提供了一个理论界限,并且我们在现实世界中以人为中心的任务中进行了广泛的实验,包括脓毒症治疗和智能辅导系统。我们的方法可靠地预测不同策略的回报,并使用标准验证方法和以人为中心的显著性测试来超越最先进的基准。
Reinforcement learning (RL) has been extensively researched for enhancing human-environment interactions in various human-centric tasks, including e-learning and healthcare. Since deploying and evaluating policies online are high-stakes in such tasks, off-policy evaluation (OPE) is crucial for inducing effective policies. In human-centric environments, however, OPE is challenging because the underlying state is often unobservable, while only aggregate rewards can be observed (students' test scores or whether a patient is released from the hospital eventually). In this work, we propose a human-centric OPE (HOPE) to handle partial observability and aggregated rewards in such environments. Specifically, we reconstruct immediate rewards from the aggregated rewards considering partial observability to estimate expected total returns. We provide a theoretical bound for the proposed method, and we have conducted extensive experiments in real-world human-centric tasks, including sepsis treatments and an intelligent tutoring system. Our approach reliably predicts the returns of different policies and outperforms state-of-the-art benchmarks using both standard validation methods and human-centric significance tests.