CIF: Small: Reinforcement Learning with Function Approximation: Convergent Algorithms and Finite-sample Analysis
CIF: Small: Reinforcement Learning with Function Approximation: Convergent Algorithms and Finite-sample Analysis
批准号:
2007783
负责人:
Shaofeng Zou
金额:
$33.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2024-09-30
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The recent success of a machine-learning technique called reinforcement learning in benchmark tasks suggests a potential revolutionary advance in practical applications, and has dramatically boosted the interest in this technique. However, common algorithms that use this approach are highly data-inefficient, leading to impressive results only on simulated systems, where an infinite amount of data can be simulated. For example, for online tasks that most humans pick up within a few minutes, reinforcement learning algorithms take much longer to reach human-level performance. A good reinforcement learning algorithm called "Rainbow deep Q-network" needs about 18 million frames of simulation data to beat human in performance for the simplest of online tasks. This amount of data corresponds to about 80 person-hours of online experience. This level of data requirements limits the application of reinforcement learning algorithms in many practical applications that only have a limited amount of data. Theoretical understanding of how much data is needed for effective reinforcement learning is still very limited. This project aims to reduce the data requirements to train reinforcement learning algorithms by developing a comprehensive methodology for reinforcement learning algorithm design and analyzing convergence rates, which will in turn motivate design of fast and stable reinforcement learning algorithms. This project will have a direct impact on various engineering and science applications, e.g., the financial market, business strategy planning, industrial automation and online advertising.This project will take a fresh perspective of using tools and concepts from both optimization and reinforcement learning. The following thrusts will be investigated in an increasing order of difficulty. 1) Linear function approximation: tools and insights will be developed to tackle challenges of non-smoothness and non-convexity in control problems. 2) General function approximation: new challenge of non-linearity will be addressed. 3) Neural function approximation: convergence to globally and/or universally optimal solutions will be investigated. In each of the three thrusts, new algorithms will be designed, and their convergence rates will be characterized. These results will be further used as guideline for parameter tuning, and to motivate design of fast and convergent algorithms.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(13)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Robust Average-Reward Markov Decision Processes
鲁棒平均奖励马尔可夫决策过程
DOI:
10.1609/aaai.v37i12.26775
发表时间:
2023
期刊:
Proceedings of the AAAI Conference on Artificial Intelligence
影响因子:
--
作者:
[Wang, Yue, Velasquez, Alvaro, Atia, George, Prater-Bennette, Ashley, Zou, Shaofeng]
通讯作者:
Zou, Shaofeng
DOI:
10.1109/iros55552.2023.10342342
发表时间:
2022-09
期刊:
2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
作者:
[Sihong He;Yue Wang;Shuo Han;Shaofeng Zou;Fei Miao]
通讯作者:
Sihong He;Yue Wang;Shuo Han;Shaofeng Zou;Fei Miao
DOI:
10.48550/arxiv.2305.10504
发表时间:
2023-05
期刊:
影响因子:
--
作者:
[Yue Wang;Alvaro Velasquez;George K. Atia;Ashley Prater-Bennette;Shaofeng Zou]
通讯作者:
Yue Wang;Alvaro Velasquez;George K. Atia;Ashley Prater-Bennette;Shaofeng Zou
DOI:
10.1109/mlsp55214.2022.9943500
发表时间:
2022-08
期刊:
2022 IEEE 32nd International Workshop on Machine Learning for Signal Processing (MLSP)
影响因子:
--
作者:
[Yudan Wang;Yue Wang;Yi Zhou;Alvaro Velasquez;Shaofeng Zou]
通讯作者:
Yudan Wang;Yue Wang;Yi Zhou;Alvaro Velasquez;Shaofeng Zou
DOI:
--
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
作者:
[Shaocong Ma;Yi Zhou;Shaofeng Zou]
通讯作者:
Shaocong Ma;Yi Zhou;Shaofeng Zou
共 10 条
CAREER: Robust Reinforcement Learning Under Model Uncertainty: Algorithms and Fundamental Limits
-
批准号:2337375
-
项目类别:Continuing Grant
-
资助金额:$52.0万
-
财政年份:2024
-
负责人:Shaofeng Zou
-
依托单位:
Collaborative Research: CIF: Medium: Emerging Directions in Robust Learning and Inference
-
批准号:2106560
-
项目类别:Continuing Grant
-
资助金额:$37.47万
-
财政年份:2021
-
负责人:Shaofeng Zou
-
依托单位:
CCSS: Collaborative Research: Quickest Threat Detection in Adversarial Sensor Networks
-
批准号:2112693
-
项目类别:Standard Grant
-
资助金额:$21.7万
-
财政年份:2021
-
负责人:Shaofeng Zou
-
依托单位:
CRII: CIF: Dynamic Network Event Detection with Time-Series Data
-
批准号:1948165
-
项目类别:Standard Grant
-
资助金额:$17.49万
-
财政年份:2020
-
负责人:Shaofeng Zou
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: