Collaborative Research: AF: Small: Parallel Reinforcement Learning with Communication and Adaptivity Constraints
Collaborative Research: AF: Small: Parallel Reinforcement Learning with Communication and Adaptivity Constraints
批准号:
2006591
负责人:
Qin Zhang
金额:
$24.22万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2023-09-30
中文摘要
近年来,强化学习的研究取得了很大的进展,并在许多实际应用中取得了成功。 然而,重复学习算法在大规模应用程序中也有数据和计算需求。 该项目将通过研究如何通过引入多个学习代理并允许它们协作收集数据和学习最佳策略来使重复学习算法可扩展的重要问题来解决这个问题。 该项目的成果将对强化学习大规模使用的许多领域产生影响,例如,多阶段临床试验、训练快速驾驶算法、众包任务、定价和不同地点商店的分类优化。 该研究成果将通过学术会议和研讨会、大学、工业实验室和在线媒体的讲座进行传播,并将整合到强化学习和大数据算法的两门前沿课程中。从技术上讲,该项目将研究如何解决学习代理在通信和自适应方面的基本限制。 特别是,本项目将研究几种协作学习模式,包括完全通信、同步通信、具有有限适应性的同步通信和并行通信,并研究以下一般性问题:(1)在并行学习模式中允许适应性的根本优势是什么;(2)基于模型和无模型的强化学习之间的并行度是否存在固有差异;(3)简化通信的影响是什么;以及(4)在强化学习中,是否可以通信有效地并行化一般算法技术? 研究团队将通过研究一系列核心问题来解决这些问题,包括多臂土匪中的最佳手臂识别和遗憾最小化,上下文土匪,有限状态马尔可夫决策过程(MDP)学习,函数近似的强化学习以及MDP中的协调探索。 通过研究这些问题,该项目将为通信高效的并行强化学习带来新的技术,观点和见解。 该项目还将对控制理论、运筹学、信息论和通信复杂性以及多智能体系统等相关研究领域产生重大影响。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Reinforcement learning has witnessed great research advancement in recent years and achieved successes in many practical applications. However, reinforcement-learning algorithms also have the reputation for being data- and computation-hungry for large-scale applications. This project will address this issue by studying the important question of how to make reinforcement-learning algorithms scalable via introducing multiple learning agents and allowing them to collect data and learn optimal strategies collaboratively. The outcomes of this project will have impacts on numerous areas where reinforcement learning is used at a scale, e.g., multi-phase clinical trials, training autonomous-driving algorithms, crowdsourcing tasks, pricing, and assortment optimization for stores at different locations. The research products will be disseminated via talks at academic conferences and workshops, universities, industrial labs, and online media, and will also be integrated in two courses on the forefront of reinforcement learning and big-data algorithms.More technically, this project will study how to address the fundamental constraints on communication and adaptivity for the learning agents. In particular, this project will investigate a handful of collaborative learning models, including full communication, synchronized communication, synchronized communication with limited adaptivity, and asynchronized communication, and study the following general questions: (1) what is the fundamental advantage of allowing adaptivity in the parallel learning model; (2) are there inherent differences on the degree of parallelism between model-based and model-free reinforcement learning; (3) what is the impact of asynchronized communication; and (4) is it possible to communication-efficiently parallelize general algorithmic techniques in reinforcement learning? The team of researchers will address these questions by studying a set of core problems, including best arm(s) identification and regret minimization in multi-armed bandits, contextual bandits, finite-state Markov decision process (MDP) learning, reinforcement learning with function approximates, and coordinated exploration in MDPs. Through studying these questions, this project will bring new techniques, perspectives, and insight to communication-efficient parallel reinforcement learning. This project will also have a significant impact on a number of related research areas such as control theory, operations research, information theory and communication complexity, and multi-agent systems.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Near-Optimal MNL Bandits Under Risk Criteria
风险标准下的近乎最优 MNL 强盗
DOI:
--
发表时间:
2021
期刊:
The Thirty-Fifth AAAI Conference on Artificial Intelligence (AAAI-21
影响因子:
--
作者:
[Xi, Guangyu, Tao, Chao, Zhou, Yuan]
通讯作者:
Zhou, Yuan
DOI:
--
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
作者:
[P. Lu;Chao Tao;Xiaojin Zhang]
通讯作者:
P. Lu;Chao Tao;Xiaojin Zhang
DOI:
10.1609/aaai.v36i7.20669
发表时间:
2020-12
期刊:
ArXiv
影响因子:
--
作者:
[Nikolai Karpov;Qin Zhang]
通讯作者:
Nikolai Karpov;Qin Zhang
DOI:
10.1109/ijcnn55064.2022.9892004
发表时间:
2022-07
期刊:
2022 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
作者:
[Boli Fang;Zhenghao Peng;Hao Sun;Qin Zhang]
通讯作者:
Boli Fang;Zhenghao Peng;Hao Sun;Qin Zhang
DOI:
--
发表时间:
2023
期刊:
Proceedings of the AAAI Conference on Artificial Intelligence (AAAI-23
影响因子:
--
作者:
[Nikolai Karpov, Qin Zhang]
通讯作者:
Nikolai Karpov, Qin Zhang
共 8 条
CAREER:Foundation of Communication-Efficient Distributed Computation and Monitoring
-
批准号:1844234
-
项目类别:Continuing Grant
-
资助金额:$49.97万
-
财政年份:2019
-
负责人:Qin Zhang
-
依托单位:
BIGDATA: Collaborative Research: F: Efficient Distributed Computation of Large-Scale Graph Problems in Epidemiology and Contagion Dynamics
-
批准号:1633215
-
项目类别:Standard Grant
-
资助金额:$53.01万
-
财政年份:2016
-
负责人:Qin Zhang
-
依托单位:
AF: Small: Redundancy exploiting algorithms for high throughput genomics
-
批准号:1619081
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2016
-
负责人:Qin Zhang
-
依托单位:
AF: Small: Efficient Algorithms for Querying Noisy Distributed/Streaming Datasets
-
批准号:1525024
-
项目类别:Standard Grant
-
资助金额:$44.43万
-
财政年份:2015
-
负责人:Qin Zhang
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: