Proto-Value Functions: A Unified Framework for Learning Task-Specific Behaviors and Task-Independent Representations
Proto-Value Functions: A Unified Framework for Learning Task-Specific Behaviors and Task-Independent Representations
批准号:
0534999
负责人:
Sridhar Mahadevan
金额:
$44.36万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-01-01 至 2009-12-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
This project addresses a longstanding puzzle in artificial intelligence (AI): how can agents transform their temporal experience into multiscale task-independent representations that can effectively guide long-term task-specific behavior? The project will investigate a nonparametric framework combining task-independent learning with task-specific learning. Algorithmically, the framework comprises of four phases. Initially, agents learn a discrete manifold representation of a given environment, which can be viewed as a topological graph representing the states reachable through single or multi-step actions. Next, the graph is analyzed using spectral clustering techniques to reveal "bottlenecks," symmetries, and other geometric invariants. In the third phase, an orthonormal set of task-independent basis functions called proto-value functions are extracted from the environment's topology: These basis functions capture large-scale geometric invariants that all value functions on the state space must adhere to. In the final phase, proto-value functions are combined with rewards to approximate task-specific value functions.The proposed framework unifies two previously disparate lines of research in AI: learning of behavior using value functions, pioneered by Arthur Samuel, and the learning of representations based on global state space analysis, pioneered by Saul Amarel. The theoretical basis for the framework draws upon links between discrete and continuous mathematics: Riemannian manifolds and the spectral theory of graphs; elliptic differential equations and abstract harmonic analysis on graphs. Specifically, the Hilbert space of smooth functions on a Riemannian manifold has a discrete spectrum based on the eigenfunctions of the Laplace-Beltrami operator. The applications of this theory to Markov decision processes will be explored, in particular the ability of Laplacian eigenfunctions or proto-value functions to both capture large-scale geometric structure and as well as approximate task-specific value functions. A novel class of algorithms termed Representation Policy Iteration (RPI) will be investigated, which interleave representation learning and behavior learning. The research thus also addresses a longstanding question not resolved in much previous work on approximation methods for solving large Markov decision processes: how can basis functions be generated automatically? The research will investigate the scalability of the proposed framework to larger problems, including both discrete factored state spaces as well as continuous state spaces. The testbeds include simulated discrete and continuous benchmark problems, simulated and real robot testbeds, and an information extraction task of maintaining the Reinforcement Learning Repository (RLR), the world's largest collection of documents and data relating to reinforcement learning.Broader impacts of this project include algorithmic and theoretical insights leading to a unified approach to learning behavior and representation, as well as applications to real-world problems such as humanoid robotics and web repository maintenance. Additionally, this project will give valuable research experience to women graduate students and to undergraduate students from local four year colleges.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Transfer Learning for Chemical Analyses from Laser-Induced Spectroscopy
-
批准号:1307179
-
项目类别:Standard Grant
-
资助金额:$15.89万
-
财政年份:2013
-
负责人:Sridhar Mahadevan
-
依托单位:
RI: Small: Reinforcement Learning by Mirror Descent
-
批准号:1216467
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2012
-
负责人:Sridhar Mahadevan
-
依托单位:
NeTS Small: Analysis and Design of Best-Effort Content-Caching Networks
-
批准号:1117764
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2011
-
负责人:Sridhar Mahadevan
-
依托单位:
Manifold Alignment of High-Dimensional Data Sets
-
批准号:1025120
-
项目类别:Standard Grant
-
资助金额:$49.99万
-
财政年份:2010
-
负责人:Sridhar Mahadevan
-
依托单位:
RI-Medium: Collaborative Research: Learning Multiscale Representations using Harmonic Analysis on Graphs
-
批准号:0803288
-
项目类别:Standard Grant
-
资助金额:$34.52万
-
财政年份:2008
-
负责人:Sridhar Mahadevan
-
依托单位:
Scaling Reinforcement Learning by Adaptive Task Selection and Linear Solution Merging
-
批准号:9896122
-
项目类别:Continuing Grant
-
资助金额:$7.85万
-
财政年份:1997
-
负责人:Sridhar Mahadevan
-
依托单位:
Scaling Reinforcement Learning by Adaptive Task Selection and Linear Solution Merging
-
批准号:9501852
-
项目类别:Continuing Grant
-
资助金额:$17.41万
-
财政年份:1995
-
负责人:Sridhar Mahadevan
-
依托单位:
Support for a Workshop on Reinforcement Learning
-
批准号:9529108
-
项目类别:Standard Grant
-
资助金额:$2.96万
-
财政年份:1995
-
负责人:Sridhar Mahadevan
-
依托单位:
国内基金
海外基金
基于时间序列间分位相依性(quantile dependence)的风险值(Value-at-Risk)预测模型研究
-
批准号:71903144
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2019
-
负责人:张申
-
依托单位: