课题基金 / 基金详情

Statistical Methods in Offline Reinforcement Learning

Statistical Methods in Offline Reinforcement Learning
离线强化学习中的统计方法
批准号:
EP/W014971/1
负责人:
Chengchun Shi
金额:
$50.76万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
强化学习(RL)关注的是智能代理如何在给定的环境中采取行动,以学习最佳策略,最大化他们收到的累积奖励。在过去的几年里,它可以说是机器学习领域最具活力的研究前沿之一。根据Google Scholar的数据,2020年已经发表了超过4万篇科学文章,其中包含“强化学习”一词。在ICML 2020(机器学习领域的重要会议)上,有100多篇关于RL的论文被接受,占被接受论文总数的10%以上。在使用RL解决各个领域的挑战性问题方面取得了重大进展,包括游戏,机器人,医疗保健,投标和自动驾驶。然而,与计算机科学相反,统计学作为一个领域,直到最近才开始在深度和广度上与RL进行接触。拟议的研究将开发统计学习方法,以解决离线RL领域的几个关键问题。我们的目标是提出RL算法,利用以前收集的数据,没有额外的在线数据收集。这项拟议的研究主要是出于医疗保健领域的应用。大多数现有的最先进的RL算法都是由在线设置(例如,视频游戏)。其在医疗保健应用中的推广仍然未知。我们还注意到,我们的解决方案将转移到其他领域(例如,机器人)。拟议的研究将考虑的一个基本问题是离线政策优化,其目标是学习最佳政策,以基于离线数据集最大化长期结果。解决这个问题至少面临两大挑战。首先,与数据易于收集或模拟的在线设置相反,许多离线应用程序中的观察次数(例如,(保健)有限。由于数据如此有限,开发具有统计效率的RL算法至关重要。拟议的研究将设计一些“价值增强”的方法,通常适用于国家的最先进的强化学习算法,以提高其统计效率。对于一个给定的初始政策计算现有的算法,我们的目标是输出一个新的政策,其预期收益收敛速度更快,实现所需的“价值增强”的属性。其次,许多离线数据集是通过聚合许多异构数据源创建的。这是医疗保健中的典型情况,其中从不同患者收集的数据轨迹可能不具有共同的分布函数。我们将研究RL中现有的迁移学习方法,并根据我们在统计学方面的专业知识开发针对医疗保健应用的新方法。拟议研究将考虑的另一个问题是政策外评估(OPE)。OPE旨在通过不同策略生成的预先收集的数据集来学习目标策略的预期回报(值)。它在医疗保健和自动驾驶等应用中至关重要,因为新策略需要在在线验证之前进行离线评估。在大多数现有的作品中,一个共同的假设是没有不可测量的混杂。然而,这一假设无法从数据中得到验证。在医疗保健应用程序生成的观察数据集中可能会违反它。此外,由于样本量有限,许多离线应用程序将受益于具有量化值估计器的不确定性的置信区间(CI)。建议的研究是关于构建一个CI的目标政策的价值存在潜在的混杂因素。此外,在各种应用中,结果分布是偏斜的和重尾的。像分位数这样的标准比平均值更合理。我们将开发方法来学习目标政策下回报的分位数曲线,并构建其相关的置信区间。
英文摘要
Reinforcement learning (RL) is concerned with how intelligent agents take actions in a given environment to learn an optimal policy that maximises the cumulative reward that they receive. It has been arguably one of the most vibrant research frontiers in machine learning over the last few years. According to Google Scholar, over 40K scientific articles have been published in 2020 with the phrase "reinforcement learning". Over 100 papers on RL were accepted for presentation at ICML 2020 (a premier conference in the machine learning area), accounting for more than 10% of the accepted papers in total. Significant progress has been made in solving challenging problems across various domains using RL, including games, robotics, healthcare, bidding and automated driving. Nevertheless statistics as a field, as opposed to computer science, has only recently begun to engage with RL both in depth and in breadth. The proposed research will develop statistical learning methodologies to address several key issues in offline RL domains. Our objective is to propose RL algorithms that utilise previously collected data, without additional online data collection. The proposed research is primarily motivated by applications in healthcare. Most of the existing state-of-the-art RL algorithms were motivated by online settings (e.g., video games). Their generalisations to applications in healthcare remain unknown. We also remark that our solutions will be transferable to other fields (e.g., robotics). A fundamental question the proposed research will consider is offline policy optimisation where the objective is to learn an optimal policy to maximise the long-term outcome based on an offline dataset. Solving this question faces at least two major challenges. First, in contrast to online settings where data are easy to collect or simulate, the number of observations in many offline applications (e.g., healthcare) is limited. With such limited data, it is critical to develop RL algorithms that are statistically efficient. The proposed research will devise some "value enhancement" methods that are generally applicable to state-of-the-art RL algorithms to improve their statistical efficiency. For a given initial policy computed by existing algorithms, we aim to output a new policy whose expected return converges at a faster rate, achieving the desired "value enhancement" property. Second, many offline datasets are created via aggregating over many heterogeneous data sources. This is typically the case in healthcare where the data trajectories collected from different patients might not have a common distribution function. We will study existing transfer learning methods in RL and develop new approaches designed for healthcare applications, based on our expertise in statistics.Another question the proposed research will consider is off-policy evaluation (OPE). OPE aims to learn a target policy's expected return (value) with a pre-collected dataset generated by a different policy. It is critical in applications from healthcare and automated driving where new policies need to be evaluated offline before online validation. A common assumption made in most of the existing works is that of no unmeasured confounding. However, this assumption is not testable from the data. It can be violated in observational datasets generated from healthcare applications. Moreover, many offline applications will benefit from having a confidence interval (CI) that quantifies the uncertainty of the value estimator, due to the limited sample size. The proposed research is concerned with constructing a CI for a target policy's value in the presence of latent confounders. In addition, in a variety of applications, the outcome distribution is skewed and heavy-tailed. Criteria such as quantiles are more sensible than the mean. We will develop methodologies to learn the quantile curve of the return under a target policy and construct its associated confidence band.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1214/22-aoas1700
发表时间: 2023-12-01
期刊: ANNALS OF APPLIED STATISTICS
影响因子: 1.8
作者: [Shi,Chengchun, Wan,Runzhe, Song,Rui]
通讯作者: Song,Rui
DOI: 10.1080/01621459.2022.2106868
发表时间: 2022-02
期刊: Journal of the American Statistical Association
影响因子: 3.7
作者: [C. Shi;S. Luo;Hongtu Zhu;R. Song]
通讯作者: C. Shi;S. Luo;Hongtu Zhu;R. Song
DOI: 10.1080/01621459.2022.2027776
发表时间: 2020-02
期刊: Journal of the American Statistical Association
影响因子: 3.7
作者: [C. Shi;Xiaoyu Wang;S. Luo;Hongtu Zhu;Jieping Ye;R. Song]
通讯作者: C. Shi;Xiaoyu Wang;S. Luo;Hongtu Zhu;Jieping Ye;R. Song
DOI: 10.1080/01621459.2023.2220169
发表时间: 2021-06
期刊: ArXiv
影响因子: --
作者: [C. Shi;Yunzhe Zhou;Lexin Li]
通讯作者: C. Shi;Yunzhe Zhou;Lexin Li
共 8 条
    国内基金
    海外基金
    Computational Methods for Analyzing Toponome Data