CAREER: Learning with Limited Feedback - Beyond Worst-case Optimality
CAREER: Learning with Limited Feedback - Beyond Worst-case Optimality
批准号:
1943607
负责人:
Haipeng Luo
金额:
$49.99万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
未结题
起止时间:
2020-03-01 至 2025-02-28
中文摘要
机器学习已经成为我们日常生活中部署的许多技术的组成部分。传统的机器学习方法首先收集数据,然后训练一个固定的模型来预测未来。然而,随着机器学习被部署在更复杂的应用中,特别是那些与人类或其他代理交互的应用,例如推荐系统,游戏代理,自动驾驶汽车等等,会出现更具挑战性的场景。在这些应用中的一个主要挑战是,学习代理通常具有来自周围环境的有限反馈,因此,利用这种有限反馈有效地学习是至关重要的。大多数现有的方法是保守的,并假设最坏的情况下的环境。该项目的重点是了解如何利用特定问题实例中表现出的特定结构,目标是开发具有强有力理论保证的自适应和高效学习算法。这个项目的成功需要在各种学科中开发新的算法技术和数学工具。通过课程开发,学生辅导,组织研讨会,并与蒙特贝洛联合学区建立合作伙伴关系,以支持建立计算机科学途径的目标,教育被融入这个项目。该项目包括三个主要方向:部分监控,强盗优化和强化学习。每个方向都在不同的维度上概括了经典的多臂强盗问题:部分监控概括了反馈模型;强盗优化概括了决策空间和目标函数;强化学习概括了从无状态到有状态的模型。每个方向包含几个主要目标:(1)对于部分监控,重点是了解如何适应数据,环境和模型;(2)对于强盗优化,重点是开发分别用于线性,凸和非凸函数学习的自适应算法;(3)对于强化学习,重点是研究在什么条件下学习变得更容易,以及如何在非平稳甚至对抗环境下学习。除了理论发展,该项目还旨在实现所有作为开源软件开发的算法,并使用基准数据集对其进行评估。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Machine learning has become an integral part of many technologies deployed in our daily lives. Traditional machine learning methods work by first collecting data and then training a fixed model for future predictions. However, much more challenging scenarios emerge as machine learning is deployed in more sophisticated applications, especially those that interact with human or other agents, such as recommender systems, game playing agents, self-driving cars, and many more. One main challenge in these applications is that the learning agent often has limited feedback from the surrounding environment, and it is thus critical to learn effectively with such limited feedback. Most existing approaches are conservative and assume worst-case environments. This project focuses on understanding how to exploit specific structures exhibited in particular problem instances, with the goal of developing more adaptive and efficient learning algorithms with strong theoretical guarantees. The success of this project requires developing new algorithmic techniques and mathematical tools in a variety of disciplines. Education is integrated into this project through curriculum development, student mentoring, organizing workshops, and developing a partnership with the Montebello Unified School District to support the goal of building Computer Science pathways. The project consists of three main directions: partial monitoring, bandit optimization, and reinforcement learning. Each direction generalizes the classic multi-armed bandit problem in a different dimension: partial monitoring generalizes the feedback model; bandit optimization generalizes the decision space and objective functions; and reinforcement learning generalizes from stateless to stateful models. Each direction contains several main objectives: (1) for partial monitoring, the focus is on understanding how to adapt to data, environments, and models; (2) for bandit optimization, the focus is on developing adaptive algorithms for learning with linear, convex, and non-convex functions respectively; (3) for reinforcement learning, the focus is on investigating under what conditions learning becomes easier, and how to learn under non-stationary or even adversarial environments. In addition to theoretical developments, the project also aims at implementing all algorithms developed as open-source software and evaluating them using benchmark datasets.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Taking a Hint: How to Leverage Loss Predictors in Contextual Bandits?
提示:如何在上下文强盗中利用损失预测器?
DOI:
--
发表时间:
2020
期刊:
Conference on Learning Theory
影响因子:
--
作者:
[Wei, Chen-Yu, Luo, Haipeng, Agarwal, Alekh]
通讯作者:
Agarwal, Alekh
DOI:
--
发表时间:
2019-10
期刊:
影响因子:
--
作者:
[Chen-Yu Wei;Mehdi Jafarnia-Jahromi;Haipeng Luo;Hiteshi Sharma;R. Jain]
通讯作者:
Chen-Yu Wei;Mehdi Jafarnia-Jahromi;Haipeng Luo;Hiteshi Sharma;R. Jain
DOI:
--
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
作者:
[Chung-Wei Lee;Haipeng Luo;Chen-Yu Wei;Mengxiao Zhang]
通讯作者:
Chung-Wei Lee;Haipeng Luo;Chen-Yu Wei;Mengxiao Zhang
DOI:
--
发表时间:
2020-02
期刊:
ArXiv
影响因子:
--
作者:
[Chung-Wei Lee;Haipeng Luo;Mengxiao Zhang]
通讯作者:
Chung-Wei Lee;Haipeng Luo;Mengxiao Zhang
Can machine learning cope with the erratic and uncertain nature of the real world?
机器学习能否应对现实世界的不稳定和不确定性?
DOI:
10.33424/futurum274
发表时间:
2022
期刊:
Futurum Careers
影响因子:
--
作者:
[Luo, Haipeng]
通讯作者:
Luo, Haipeng
共 8 条
CRII:RI: Adaptive and Practical Algorithms for Personalization
-
批准号:1755781
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2018
-
负责人:Haipeng Luo
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:吉建娇
-
依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
-
批准号:62003314
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:沈剑
-
依托单位:
集成上下文张量分解的e-learning资源推荐方法研究
-
批准号:61902016
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2019
-
负责人:万珊珊
-
依托单位:
具有时序迁移能力的Spiking-Transfer learning (脉冲-迁移学习)方法研究
-
批准号:61806040
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2018
-
负责人:解修蕊
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于有向超图的大型个性化e-learning学习过程模型的自动生成与优化
-
批准号:61572533
-
项目类别:面上项目
-
资助金额:66.0万元
-
批准年份:2015
-
负责人:孙雪冬
-
依托单位:
E-Learning中学习者情感补偿方法的研究
-
批准号:61402392
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2014
-
负责人:秦继伟
-
依托单位: