课题基金 / 基金详情

Proto-Value Functions: A Unified Framework for Learning Task-Specific Behaviors and Task-Independent Representations

Proto-Value Functions: A Unified Framework for Learning Task-Specific Behaviors and Task-Independent Representations
原始价值函数:学习任务特定行为和任务无关表示的统一框架
批准号:
0534999
负责人:
Sridhar Mahadevan
金额:
$44.36万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-01-01 至 2009-12-31

项目摘要

项目成果

Sridhar Mahadevan的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
This project addresses a longstanding puzzle in artificial intelligence (AI): how can agents transform their temporal experience into multiscale task-independent representations that can effectively guide long-term task-specific behavior? The project will investigate a nonparametric framework combining task-independent learning with task-specific learning. Algorithmically, the framework comprises of four phases. Initially, agents learn a discrete manifold representation of a given environment, which can be viewed as a topological graph representing the states reachable through single or multi-step actions. Next, the graph is analyzed using spectral clustering techniques to reveal "bottlenecks," symmetries, and other geometric invariants. In the third phase, an orthonormal set of task-independent basis functions called proto-value functions are extracted from the environment's topology: These basis functions capture large-scale geometric invariants that all value functions on the state space must adhere to. In the final phase, proto-value functions are combined with rewards to approximate task-specific value functions.The proposed framework unifies two previously disparate lines of research in AI: learning of behavior using value functions, pioneered by Arthur Samuel, and the learning of representations based on global state space analysis, pioneered by Saul Amarel. The theoretical basis for the framework draws upon links between discrete and continuous mathematics: Riemannian manifolds and the spectral theory of graphs; elliptic differential equations and abstract harmonic analysis on graphs. Specifically, the Hilbert space of smooth functions on a Riemannian manifold has a discrete spectrum based on the eigenfunctions of the Laplace-Beltrami operator. The applications of this theory to Markov decision processes will be explored, in particular the ability of Laplacian eigenfunctions or proto-value functions to both capture large-scale geometric structure and as well as approximate task-specific value functions. A novel class of algorithms termed Representation Policy Iteration (RPI) will be investigated, which interleave representation learning and behavior learning. The research thus also addresses a longstanding question not resolved in much previous work on approximation methods for solving large Markov decision processes: how can basis functions be generated automatically? The research will investigate the scalability of the proposed framework to larger problems, including both discrete factored state spaces as well as continuous state spaces. The testbeds include simulated discrete and continuous benchmark problems, simulated and real robot testbeds, and an information extraction task of maintaining the Reinforcement Learning Repository (RLR), the world's largest collection of documents and data relating to reinforcement learning.Broader impacts of this project include algorithmic and theoretical insights leading to a unified approach to learning behavior and representation, as well as applications to real-world problems such as humanoid robotics and web repository maintenance. Additionally, this project will give valuable research experience to women graduate students and to undergraduate students from local four year colleges.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Transfer Learning for Chemical Analyses from Laser-Induced Spectroscopy
  • 批准号:
    1307179
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.89万
  • 财政年份:
    2013
  • 负责人:
    Sridhar Mahadevan
  • 依托单位:
RI: Small: Reinforcement Learning by Mirror Descent
  • 批准号:
    1216467
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2012
  • 负责人:
    Sridhar Mahadevan
  • 依托单位:
NeTS Small: Analysis and Design of Best-Effort Content-Caching Networks
  • 批准号:
    1117764
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2011
  • 负责人:
    Sridhar Mahadevan
  • 依托单位:
Manifold Alignment of High-Dimensional Data Sets
  • 批准号:
    1025120
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.99万
  • 财政年份:
    2010
  • 负责人:
    Sridhar Mahadevan
  • 依托单位:
国内基金
海外基金
基于时间序列间分位相依性(quantile dependence)的风险值(Value-at-Risk)预测模型研究
  • 批准号:
    71903144
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    17.0万元
  • 批准年份:
    2019
  • 负责人:
    张申
  • 依托单位: