课题基金 / 基金详情

Proto-Value Functions: A Unified Framework for Learning Task-Specific Behaviors and Task-Independent Representations

Proto-Value Functions: A Unified Framework for Learning Task-Specific Behaviors and Task-Independent Representations
原始价值函数:学习任务特定行为和任务无关表示的统一框架
批准号:
0534999
负责人:
Sridhar Mahadevan
金额:
$44.36万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-01-01 至 2009-12-31

项目摘要

项目成果

Sridhar Mahadevan的其他基金

相似基金

相关文献

中文摘要
翻译
这个项目解决了人工智能(AI)中一个长期存在的难题:代理如何将他们的时间经验转换为多尺度的任务独立表示,从而有效地指导长期的任务特定行为?该项目将研究一个结合了任务独立学习和任务特定学习的非参数框架。在算法上,该框架由四个阶段组成。最初,代理学习给定环境的离散流形表示,它可以被视为表示通过单步或多步操作可到达的状态的拓扑图。接下来,使用谱聚类技术对图进行分析,以揭示“瓶颈”、对称性和其他几何不变量。在第三阶段,从环境的拓扑结构中提取一组与任务无关的正交基函数,称为原值函数:这些基函数捕获状态空间上所有值函数必须遵守的大规模几何不变量。在最后阶段,将原始价值函数与奖励相结合来近似任务特定的价值函数。该框架统一了人工智能中两个以前截然不同的研究方向:由Arthur Samuel开创的使用价值函数学习行为,以及由Saul Amarel开创的基于全局状态空间分析的表征学习。该框架的理论基础借鉴了离散数学和连续数学之间的联系:黎曼流形和图的谱理论;椭圆型微分方程和图上的抽象调和分析。具体地说,黎曼流形上光滑函数的Hilbert空间具有基于Laplace-Beltrami算子本征函数的离散谱。我们将探讨这一理论在马尔可夫决策过程中的应用,特别是拉普拉斯特征函数或原值函数捕捉大规模几何结构以及近似任务特定值函数的能力。研究了一类新的交叉表示学习和行为学习的算法,称为表示策略迭代算法(RPI)。因此,这项研究还解决了一个长期存在的问题,即如何自动生成基函数,这一问题在以前关于求解大型马尔可夫决策过程的近似方法的工作中没有得到解决。这项研究将考察所提出的框架对更大问题的可扩展性,包括离散因式分解状态空间和连续状态空间。试验台包括模拟的离散和连续基准问题,模拟和真实的机器人试验台,以及维护强化学习存储库(RLR)的信息提取任务,RLR是世界上与强化学习相关的最大文档和数据集合。该项目的广泛影响包括算法和理论见解,导致学习行为和表示的统一方法,以及应用于现实世界的问题,如人形机器人和Web存储库维护。此外,该项目将为当地四年制大学的女研究生和本科生提供宝贵的研究经验。
英文摘要
This project addresses a longstanding puzzle in artificial intelligence (AI): how can agents transform their temporal experience into multiscale task-independent representations that can effectively guide long-term task-specific behavior? The project will investigate a nonparametric framework combining task-independent learning with task-specific learning. Algorithmically, the framework comprises of four phases. Initially, agents learn a discrete manifold representation of a given environment, which can be viewed as a topological graph representing the states reachable through single or multi-step actions. Next, the graph is analyzed using spectral clustering techniques to reveal "bottlenecks," symmetries, and other geometric invariants. In the third phase, an orthonormal set of task-independent basis functions called proto-value functions are extracted from the environment's topology: These basis functions capture large-scale geometric invariants that all value functions on the state space must adhere to. In the final phase, proto-value functions are combined with rewards to approximate task-specific value functions.The proposed framework unifies two previously disparate lines of research in AI: learning of behavior using value functions, pioneered by Arthur Samuel, and the learning of representations based on global state space analysis, pioneered by Saul Amarel. The theoretical basis for the framework draws upon links between discrete and continuous mathematics: Riemannian manifolds and the spectral theory of graphs; elliptic differential equations and abstract harmonic analysis on graphs. Specifically, the Hilbert space of smooth functions on a Riemannian manifold has a discrete spectrum based on the eigenfunctions of the Laplace-Beltrami operator. The applications of this theory to Markov decision processes will be explored, in particular the ability of Laplacian eigenfunctions or proto-value functions to both capture large-scale geometric structure and as well as approximate task-specific value functions. A novel class of algorithms termed Representation Policy Iteration (RPI) will be investigated, which interleave representation learning and behavior learning. The research thus also addresses a longstanding question not resolved in much previous work on approximation methods for solving large Markov decision processes: how can basis functions be generated automatically? The research will investigate the scalability of the proposed framework to larger problems, including both discrete factored state spaces as well as continuous state spaces. The testbeds include simulated discrete and continuous benchmark problems, simulated and real robot testbeds, and an information extraction task of maintaining the Reinforcement Learning Repository (RLR), the world's largest collection of documents and data relating to reinforcement learning.Broader impacts of this project include algorithmic and theoretical insights leading to a unified approach to learning behavior and representation, as well as applications to real-world problems such as humanoid robotics and web repository maintenance. Additionally, this project will give valuable research experience to women graduate students and to undergraduate students from local four year colleges.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Transfer Learning for Chemical Analyses from Laser-Induced Spectroscopy
  • 批准号:
    1307179
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.89万
  • 财政年份:
    2013
  • 负责人:
    Sridhar Mahadevan
  • 依托单位:
RI: Small: Reinforcement Learning by Mirror Descent
  • 批准号:
    1216467
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2012
  • 负责人:
    Sridhar Mahadevan
  • 依托单位:
NeTS Small: Analysis and Design of Best-Effort Content-Caching Networks
  • 批准号:
    1117764
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2011
  • 负责人:
    Sridhar Mahadevan
  • 依托单位:
Manifold Alignment of High-Dimensional Data Sets
  • 批准号:
    1025120
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.99万
  • 财政年份:
    2010
  • 负责人:
    Sridhar Mahadevan
  • 依托单位:
国内基金
海外基金
基于时间序列间分位相依性(quantile dependence)的风险值(Value-at-Risk)预测模型研究
  • 批准号:
    71903144
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    17.0万元
  • 批准年份:
    2019
  • 负责人:
    张申
  • 依托单位: