课题基金 / 基金详情

Collaborative Research: Towards the Foundation of Approximate Sampling-Based Exploration in Sequential Decision Making

Collaborative Research: Towards the Foundation of Approximate Sampling-Based Exploration in Sequential Decision Making
协作研究:为顺序决策中基于近似采样的探索奠定基础
批准号:
2323113
负责人:
Quanquan Gu
金额:
$30.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2026-09-30

项目摘要

项目成果

Quanquan Gu的其他基金

相似基金

相关文献

中文摘要
翻译
顺序决策问题,如强盗和强化学习,在各种人工智能应用中发挥着至关重要的作用,包括推荐系统、机器人、游戏和个性化医疗保健。主要的挑战在于找到最优的探索策略,在选择具有最佳性能的行动和选择具有高不确定性的行动之间取得平衡。然而,现有的勘探策略严重依赖于具体情况,需要事先了解奖励分布、函数近似和手头的任务。这就造成了计算障碍,阻碍了现实世界的适用性。本项目旨在为利用基于近似采样的技术统一不同序列决策问题的勘探策略奠定理论基础。目标是在基于近似采样的统一算法框架下,开发适用于各种学习问题的高效且可证明的算法。本项目也为研究生提供了研究训练的机会。该项目包括三个任务。任务一的重点是为背景土匪问题开发基于近似采样的快速勘探策略,并提供理论保证。任务二涉及实现和推广这些探索算法到更复杂的顺序决策应用,利用深度神经网络。任务三旨在为强化学习问题建立高效且可证明有效的探索策略。这些进步将转化为各种强盗和强化学习应用程序的可访问工具,提供可验证的保证。这个项目产生的开源软件和课程材料将会公开,使研究、教育和整个社会受益。该奖项由数学科学部颁发,由国家科学基金会高级网络基础设施办公室联合支持。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Sequential decision-making problems, such as bandits and reinforcement learning, play a crucial role in various AI applications, including recommendation systems, robotics, games, and personalized healthcare. The main challenge lies in finding the optimal exploration strategy that strikes a balance between choosing actions with the best performance and choosing actions with high uncertainties. However, existing exploration strategies heavily depend on specific cases, requiring prior knowledge of reward distribution, function approximation, and the task at hand. This creates computational obstacles and hampers real-world applicability. This project aims to establish a theoretical foundation for using approximate sampling-based techniques to unify exploration strategies across different sequential decision problems. The goal is to develop efficient and provable algorithms applicable to diverse learning problems under a unified algorithmic framework based on approximate sampling. This project also provides research training opportunities for graduate students. The project consists of three tasks. Task one focuses on developing fast approximate sampling-based exploration strategies for contextual bandit problems, accompanied by theoretical guarantees. Task two involves implementing and generalizing these exploration algorithms to more complex sequential decision-making applications, leveraging deep neural networks. Task three aims to establish efficient and provably effective exploration strategies for reinforcement learning problems. These advancements will be translated into accessible tools for various bandit and reinforcement learning applications, providing verifiable guarantees. The open-source software and course materials resulting from this project will be made publicly available, benefiting research, education, and society at large.This award by the Division of Mathematical Sciences is jointly supported by the NSF Office of Advanced Cyberinfrastructure.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CPS: Medium: Collaborative Research: Provably Safe and Robust Multi-Agent Reinforcement Learning with Applications in Urban Air Mobility
III: Small: Towards the Foundations of Training Deep Neural Networks: New Theory and Algorithms
CIF: Small: Collaborative Research: Rank Aggregation with Heterogeneous Information Sources: Efficient Algorithms and Fundamental Limits
III: Small: Collaborative Research: High-Dimensional Machine Learning Methods for Personalized Cancer Genomics
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)