课题基金 / 基金详情

Interactive reinforcement learning for adaptive experimental design

Interactive reinforcement learning for adaptive experimental design
用于自适应实验设计的交互式强化学习
批准号:
RGPIN-2020-06933
负责人:
Durand, Audrey
金额:
$1.75万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2020
资助国家:
加拿大
项目状态:
已结题
起止时间:
2020-01-01 至 2021-12-31

项目摘要

项目成果

Durand, Audrey的其他基金

相似基金

相关文献

中文摘要
翻译
与研究现有数据的静态集合的监督和无监督学习算法不同,强化学习(RL)算法选择动作来改变其环境,并使用这些动作的影响来随着时间的推移逐步改进其选择。RL设置的最简单的例子是强盗问题,其中代理人面临一个固定的动作集之间的单一选择,并在一个时间步长内从环境中接收反馈。最初的动机是自适应实验设计,其中一个目的是比较几个行动进行实验,在最大限度地提供信息和有效的方式,强盗算法已被深入研究在过去的十年。除了理论发现,这些算法也已成功地在适应性实验中实现,例如用于优化小鼠试验中的癌症治疗,用于调整显微镜成像系统,用于调整超参数,以及用于设计用户界面。然而,许多现实世界的应用程序的特点是动态的,没有捕捉到现有的RL设置。更糟糕的是,这些动态有时与当前算法所需的假设相矛盾。这些情况导致算法具有无效的理论保证,有可能产生不受欢迎的,潜在的危险行为。因此,在实践中利用交互式RL算法的力量需要针对这些特定条件的方法。 拟议的研究计划旨在将RL带到适应性实验设计的真实的世界。我们的目标是开发出符合预期的算法,并在这些现实环境下保持理论保证。通过与其他领域的研究人员合作,我们将部署这些策略,并衡量它们对真实的应用程序的影响。 该研究计划围绕以下目标进行阐述: 1)提出灵活的交互式RL设置,包括最先进的自适应实验设计框架; 2)研究现实的系统动力学,并在这些约束条件下引入理论基础的学习算法; 3)描述学习算法中出现的不良行为,并制定防范策略; 4)在实际应用中部署算法,以展示其潜力并影响应用领域。 拟议的研究方向具有很高的影响潜力,因为它将导致为真实的应用案例开发的策略。尽管该计划的重点是强盗子字段,产生的知识和算法将构成理解和设计顺序RL算法的基础,这仍然基本上仍然局限于仿真环境。最后,这些应用程序将构成概念证明,并将产生安全和有效的RL部署指南,支持进一步的研究,使RL更接近该领域。
英文摘要
Unlike algorithms in supervised and unsupervised learning, which study static collections of existing data, algorithms in reinforcement learning (RL) choose actions to transform their environments and use the effects of those actions to improve their choices incrementally over time. The simplest instance of an RL setting is the bandit problem, in which an agent faces a single choice between a fixed set of actions and receives feedback from its environment within one time-step. Initially motivated by adaptive experimental designs, where one aims to compare several actions by conducting experiments in a maximally informative and efficient way, bandit algorithms have been studied intensively throughout the last decade. Beyond theoretical findings, these algorithms have also been successfully implemented in adaptive experiments, for example for optimizing cancer treatments in mice trials, for tuning microscopy imaging systems, for adjusting hyperparameters, and for designing user interfaces. However, many real-world applications are characterized by dynamics that are not captured by existing RL settings. Even worse, these dynamics sometimes contradict assumptions required by current algorithms. These cases result in algorithms with voided theoretical guarantees, at risk of producing undesirable, potentially dangerous, behaviours. Therefore, leveraging the power of interactive RL algorithms in practice requires approaches intended for these specific conditions. The proposed research program aims to bring RL to the real world of adaptive experimental design. We aim at developing algorithms that perform as expected and with theoretical guarantees that hold under these realistic environments. Through collaborations with researchers in other fields, we will deploy those strategies and measure their impact on real applications. The research program is articulated around the following objectives: 1) Propose flexible interactive RL settings that encompass state-of-the-art adaptive experimental design frameworks; 2) Investigate realistic system dynamics and introduce theoretically grounded algorithms for learning under these constraints; 3) Characterize the arising of undesirable behaviours in learning algorithms and develop strategies to guard against it; 4) Deploy algorithms in real-world applications to showcase their potential and impact the application domain. The proposed research direction has a high-impact potential as it will result in strategies developed for real application cases. Even though the program is focused on the bandit subfield, resulting knowledge and algorithms will constitute a basis for understanding and designing sequential RL algorithms, which still essentially remain confined to simulation environments. Finally, the applications will constitute proofs of concept and will result in deployment guidelines for safe and impactful RL, supporting further research that will bring RL closer to the field.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Interactive reinforcement learning for adaptive experimental design
  • 批准号:
    RGPIN-2020-06933
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.75万
  • 财政年份:
    2022
  • 负责人:
    Durand, Audrey
  • 依托单位:
Interactive reinforcement learning for adaptive experimental design
  • 批准号:
    RGPIN-2020-06933
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.75万
  • 财政年份:
    2021
  • 负责人:
    Durand, Audrey
  • 依托单位:
Interactive reinforcement learning for adaptive experimental design
  • 批准号:
    DGECR-2020-00327
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2020
  • 负责人:
    Durand, Audrey
  • 依托单位:
Self-management of distributed resources in wireless sensor networks
  • 批准号:
    443420-2013
  • 项目类别:
    Postgraduate Scholarships - Doctoral
  • 资助金额:
    $1.53万
  • 财政年份:
    2014
  • 负责人:
    Durand, Audrey
  • 依托单位:
国内基金
海外基金
海桑属杂种区强化(Reinforcement)的检验与遗传基础研究
  • 批准号:
    30800060
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    23.0万元
  • 批准年份:
    2008
  • 负责人:
    周仁超
  • 依托单位: