CIF: Small: Compression Schemes for Communication Constrained Bandit and Reinforcement Learning
CIF: Small: Compression Schemes for Communication Constrained Bandit and Reinforcement Learning
批准号:
2221871
负责人:
Lin Yang
金额:
$60.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-10-01 至 2025-09-30
中文摘要
主动学习和在线学习是机器学习的范例,其中计算机在接受环境反馈的同时学习做出复杂的决策。例如,无人机可以自己学习飞行,或者汽车可以通过反复试验来学习驾驶。最近,这些学习范式得到了广泛应用,并在游戏玩法或机器人控制等任务中取得了显著的成功。随着计算设备变得越来越小,功耗越来越低,新的分布式学习框架开始出现。这些框架包含低能力的学习代理(如手机、无人驾驶车辆或无人机),它们相距很远,但通过(无线)网络相互通信来集体执行学习。然而,现有的通信方法是为高性能计算机设计的,消耗了太多的功率和网络带宽,因此会成为学习的瓶颈。该项目旨在通过提供新颖的技术来解决这一问题,该技术可以有效地压缩数据以进行通信,同时保留学习能力。本项目开发的技术将通过提高交流效率,推动分布式在线/主动学习的发展。该项目的总体目标是建立有效的压缩方案,支持有效的主动/在线学习,例如通信受限网络上的强盗和强化学习。在这些学习环境中,学习者的目标是根据经验为下一步做出正确的决定;该项目将探索支持这一目标的基本界限和有效算法,同时通过压缩只保留决策所需信息的方式,最大限度地减少通信位数。换句话说,这个项目旨在探索在活动/在线环境中压缩和可学习性之间的基本权衡。在有希望的初步工作的基础上,研究人员将研究从最基本的多臂强盗设置到更复杂的强化学习设置的问题,并考虑集中和分散的网络拓扑结构。更具体地说,研究人员提出了以下问题的压缩方案和基本理论界限:(1)多武装强盗问题中的奖励;(2)上下文强盗问题的上下文向量;(3)马尔可夫决策问题的状态-行为特征和模型。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Active learning and online learning are machine-learning paradigms in which computers learn to make complex decisions while receiving feedback from an environment. For instance, a drone may learn to fly by itself, or a car may learn to drive by trial and error. Recently, these learning paradigms have been widely applied and have achieved phenomenal successes with human-level performance in tasks like gameplay or robot control. As computing devices become smaller and less power-consuming, new distributed learning frameworks start to emerge. These frameworks contain low-capability learning agents (such as cell phones, unmanned vehicles, or drones) that are far apart but perform learning collectively by communicating with each other through (wireless) networks. However, existing communication approaches would become bottlenecks for learning since they were designed for high-power computers and consume too much power and network bandwidth. This project aims to address this issue by providing novel techniques that efficiently compress data to be communicated while preserving the learning ability. The techniques developed in this project will advance the state-of-the-art in distributed online/active learning by improving communication efficiencies. The overarching goal of this project is to establish efficient compression schemes that support effective active/online learning, such as bandit and reinforcement learning over communication-constrained networks. In these learning environments, a learner aims to make a good decision for the next steps based on experience; this project will explore fundamental bounds and efficient algorithms that support this goal while minimizing the number of bits communicated - by compressing in a way that only retains the necessary information for decision making. In other words, this project aims to explore the fundamental trade-off between compression and learnability in active/online environments. Building on promising preliminary work, the investigators will study problems ranging from the most basic multi-arm bandit setting to more complex reinforcement learning settings and consider both centralized and decentralized network topologies. More specifically, the investigators propose compression schemes and fundamental theoretical bounds for (1) rewards in multi-armed bandit problems, (2) context vectors for contextual bandit problems, and (3) state-action features and models for Markov decision problems.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Near-Optimal Sample Complexity Bounds for Constrained MDPs
受限 MDP 的近乎最优样本复杂度界限
DOI:
--
发表时间:
2022
期刊:
Advances in neural information processing systems
影响因子:
--
作者:
[Vaswani, Sharan, Yang, Lin, Szepesvári, Csaba]
通讯作者:
Szepesvári, Csaba
DOI:
10.48550/arxiv.2304.08944
发表时间:
2023-04
期刊:
ArXiv
影响因子:
--
作者:
[Dingwen Kong;Lin F. Yang]
通讯作者:
Dingwen Kong;Lin F. Yang
PROVABLY EFFICIENT LIFELONG REINFORCEMENT LEARNING WITH LINEAR REPRESENTATION
具有线性表示的可证明有效的终身强化学习
DOI:
--
发表时间:
2023
期刊:
ICLR
影响因子:
--
作者:
[Amani, Sanae, Yang, Lin, Cheng, Ching-An]
通讯作者:
Cheng, Ching-An
DOI:
10.48550/arxiv.2306.09554
发表时间:
2023-06
期刊:
ArXiv
影响因子:
--
作者:
[Yunfan Li-;Yiran Wang-;Y. Cheng;Lin F. Yang]
通讯作者:
Yunfan Li-;Yiran Wang-;Y. Cheng;Lin F. Yang
Horizon-Free Learning for Markov Decision Processes and Games: Stochastically Bounded Rewards and Improved Bounds
马尔可夫决策过程和博弈的无地平线学习:随机有界奖励和改进界限
DOI:
--
发表时间:
2023
期刊:
Proceedings of Machine Learning Research
影响因子:
--
作者:
[Li, Shengshi, Yang, Lin]
通讯作者:
Yang, Lin
共 9 条
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: