课题基金 / 基金详情

Offline Statistical Reinforcement Learning with Applications in Precision Health

Offline Statistical Reinforcement Learning with Applications in Precision Health
离线统计强化学习在精准健康中的应用
批准号:
2113637
负责人:
Wenbin Lu
金额:
$20.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-08-15 至 2024-07-31

项目摘要

项目成果

Wenbin Lu的其他基金

相似基金

相关文献

中文摘要
翻译
精准医学旨在根据每个患者的个体特征量身定制医疗,以达到更好的患者治疗效果。作为一个更广泛的概念,包括精准医学,精准健康包括每个人都可以自己采取的方法来保护自己的健康,以及公共卫生可以采取的步骤。强化学习(RL)是一种强大的技术,它允许代理在给定的环境中学习和采取行动,以最大化代理收到的累积奖励。对开发新的精确健康统计RL方法的兴趣正在出现。这项工作的潜在影响可以概括为以下四个目标。首先,该项目对半参数推理和强化学习领域都有贡献。理论结果包括非渐近分布,风险界限与新的经验过程技术工具。这些结果对于研究强化学习的半参数推理具有重要的基础意义和普遍适用性。其次,基于电子病历(EMR)数据分析的临床发现将在解决重要的临床问题上取得重大进展,为患者提供治疗建议。第三,虽然EMR和移动健康(mHealth)数据是这个项目的主要应用,但开发的方法是通用的,足以适用于各种数据源,包括临床数据和经济数据。所开发的方法有望大大增强医学、科学和工程界对具有人口异质性的大规模数据的获取和分析。第四,科研与教育的融合是该项目的一个关键方面。PI将开发新的课程,并改进现有的RL和半参数推理课程,将培养研究生,并将通过培训高中教师和学生,扩大到K-12教育水平。尽管强化学习在游戏和机器人等领域取得了巨大的影响,但由于现实世界的重大挑战,在精确健康领域直接部署强化学习算法可能成本高昂,风险很大,甚至是不可行的。本提案的主要目标是开发新的统计离线RL方法,通过开发灵活高效的非政策学习和稳健高效的非政策评估方法来处理现实世界的挑战。RL中的off-policy学习是指从可能不同的策略中收集样本,找到使价值最大化的最佳目标策略的问题。在目标1中,我们将开发一个有效的优势学习框架,以便有效地使用预先收集的数据进行政策优化。在目标2中,我们考虑了政策外评估的问题,其目标是使用在可能不同的政策下收集的数据来学习目标政策下的价值。关于在非政策设置下估计给定政策下的价值的文献越来越多。然而,对于统计推断,如假设检验和值的置信区间(ci),已经考虑了非常有限的工作,这是本目标的重点。在目标3中,我们将讨论基于目标1和2中提出的方法,用政策学习和政策评估来分析EMR和移动健康数据的计划。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Precision medicine seeks to tailor medical treatment to the individual characteristics of each patient to achieve the goal of better patient outcomes. As a broader conceptualization that includes precision medicine, precision health involves approaches that everyone can do on their own to protect their health as well as steps that public health can take. Reinforcement Learning (RL) is a powerful technique that allows an agent to learn and take actions in a given environment in order to maximize the cumulative reward that the agent receives. The interest in developing new statistical RL methods for precision health is emerging. The potential impacts of this work can be summarized in the following four goals. First, the project contributes to both the fields of semiparametric inference and RL. The theoretical results include non-asymptotic distribution, risk bounds with novel empirical process technical tools. These results will be fundamentally important and generally applicable for studying semiparametric inference empowered by RL. Second, the clinical findings based on analyzing the electronic medical record (EMR) data will lead to major progress in addressing important clinical questions on the treatment recommendations for patients. Third, although EMR and mobile health (mHealth) data are the main applications of this project, the developed methods are general enough to apply to a variety of data sources including clinical data and economic data. The developed methods are expected to greatly enhance the acquisition and analysis of large-scale data with population heterogeneity for medical, scientific and engineering communities. Fourth, the integration of research and education is a key aspect of this project. The PI will develop new courses and improve existing courses on RL and semiparametric inference, will train graduate students, and will reach out to the K-12 education levels by training high school teachers and students.Despite the tremendous impacts that RL has achieved in areas such as games and robots, a direct deployment of RL algorithms in precision health can be costly, risky or even infeasible, due to significant real-world challenges. The main objective of this proposal is to develop new statistical offline RL methods to handle real-world challenges by developing flexible and efficient off-policy learning and robust and efficient off-policy evaluation methods. The off-policy learning in RL refers to the problem of finding the best target policy that maximizes the value, given samples collected from a possibly different policy. In Aim 1, we will develop an efficient advantage learning framework in order to efficiently use pre-collected data for policy optimization. In Aim 2, we consider the problem of off-policy evaluation where the objective is to learn the value under a target policy with data collected under a possibly different policy. There is a growing literature on estimating the value under a given policy in off-policy settings. However, very limited work have been considered regarding statistical inference such as hypothesis testing and confidence intervals (CIs) of the value, which is the focus of this aim. In Aim 3, we will discuss the plan of analyzing EMR and mHealth data with policy learning and policy evaluation based on the proposed methods in Aims 1 and 2.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2022-02
期刊: ArXiv
影响因子: --
作者: [Runzhe Wan;B. Kveton;Rui Song]
通讯作者: Runzhe Wan;B. Kveton;Rui Song
Semiparametric Models, Methodologies and Related Theory for Analysis of Censored Survival Data
  • 批准号:
    0504269
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $4.98万
  • 财政年份:
    2005
  • 负责人:
    Wenbin Lu
  • 依托单位:
海外基金