Collaborative Research: AF: Small: Parallel Reinforcement Learning with Communication and Adaptivity Constraints
Collaborative Research: AF: Small: Parallel Reinforcement Learning with Communication and Adaptivity Constraints
批准号:
2006526
负责人:
Yuan Zhou
金额:
$25.77万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2024-09-30
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Reinforcement learning has witnessed great research advancement in recent years and achieved successes in many practical applications. However, reinforcement-learning algorithms also have the reputation for being data- and computation-hungry for large-scale applications. This project will address this issue by studying the important question of how to make reinforcement-learning algorithms scalable via introducing multiple learning agents and allowing them to collect data and learn optimal strategies collaboratively. The outcomes of this project will have impacts on numerous areas where reinforcement learning is used at a scale, e.g., multi-phase clinical trials, training autonomous-driving algorithms, crowdsourcing tasks, pricing, and assortment optimization for stores at different locations. The research products will be disseminated via talks at academic conferences and workshops, universities, industrial labs, and online media, and will also be integrated in two courses on the forefront of reinforcement learning and big-data algorithms.More technically, this project will study how to address the fundamental constraints on communication and adaptivity for the learning agents. In particular, this project will investigate a handful of collaborative learning models, including full communication, synchronized communication, synchronized communication with limited adaptivity, and asynchronized communication, and study the following general questions: (1) what is the fundamental advantage of allowing adaptivity in the parallel learning model; (2) are there inherent differences on the degree of parallelism between model-based and model-free reinforcement learning; (3) what is the impact of asynchronized communication; and (4) is it possible to communication-efficiently parallelize general algorithmic techniques in reinforcement learning? The team of researchers will address these questions by studying a set of core problems, including best arm(s) identification and regret minimization in multi-armed bandits, contextual bandits, finite-state Markov decision process (MDP) learning, reinforcement learning with function approximates, and coordinated exploration in MDPs. Through studying these questions, this project will bring new techniques, perspectives, and insight to communication-efficient parallel reinforcement learning. This project will also have a significant impact on a number of related research areas such as control theory, operations research, information theory and communication complexity, and multi-agent systems.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(12)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2018-10
期刊:
J. Mach. Learn. Res.
影响因子:
--
作者:
[Xi Chen;Yining Wang;Yuanshuo Zhou]
通讯作者:
Xi Chen;Yining Wang;Yuanshuo Zhou
Collaborative Top Distribution Identifications with Limited Interaction (Extended Abstract)
有限交互的协作顶级分布识别(扩展摘要)
DOI:
10.1109/focs46700.2020.00024
发表时间:
2020
期刊:
2020 IEEE 61st Annual Symposium on Foundations of Computer Science (FOCS
影响因子:
--
作者:
[Karpov, Nikolai, Zhang, Qin, Zhou, Yuan]
通讯作者:
Zhou, Yuan
Linear bandits with limited adaptivity and learning distributional optimal design
具有有限适应性和学习分布优化设计的线性老虎机
DOI:
10.1145/3406325.3451004
发表时间:
2021
期刊:
STOC 2021: Proceedings of the 53rd Annual ACM SIGACT Symposium on Theory of Computing
影响因子:
--
作者:
[Ruan, Yufei, Yang, Jiaqi, Zhou, Yuan]
通讯作者:
Zhou, Yuan
Dynamic Assortment Planning Under Nested Logit Models
嵌套 Logit 模型下的动态分类规划
DOI:
10.1111/poms.13258
发表时间:
2021
期刊:
Production and Operations Management
影响因子:
5
作者:
[Chen, Xi, Shi, Chao, Wang, Yining, Zhou, Yuan]
通讯作者:
Zhou, Yuan
Dynamic Pricing and Inventory Control with Fixed Ordering Cost and Incomplete Demand Information
固定订购成本和不完整需求信息的动态定价和库存控制
DOI:
10.1287/mnsc.2021.4171
发表时间:
2021
期刊:
Management Science
影响因子:
5.4
作者:
[Chen, Boxiao, Simchi-Levi, David, Wang, Yining, Zhou, Yuan]
通讯作者:
Zhou, Yuan
共 12 条
Collaborative Research: Next-Generation Cutting Planes: Compression, Automation, Diversity, and Computer-Assisted Mathematics
-
批准号:2012429
-
项目类别:Standard Grant
-
资助金额:$17.98万
-
财政年份:2020
-
负责人:Yuan Zhou
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: