课题基金 / 基金详情

CRII: RI: RUI: Performance guarantees for online apprenticeship learning with unknown features

CRII: RI: RUI: Performance guarantees for online apprenticeship learning with unknown features
CRII:RI:RUI:具有未知特征的在线学徒学习的性能保证
批准号:
1850149
负责人:
Kenneth Bogert
金额:
$15.83万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-05-01 至 2022-04-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
The surge of interest in robots that can be trained to perform in industries such as manufacturing and healthcare increases the need for improved learning methods. In one such method, apprenticeship learning, a robot learns to perform a task by watching an expert. This project's goals are to decrease the time required to set up the robot for learning and to offer college students hands-on robotic research activities. Most related work describes techniques that require the robot's programmer to identify features of the task. This project will reduce the programmer's setup work by using automatically generated features. A new method ensures the accuracy of the robotic learner by determining the number of observations required of the expert. Maximum Causal Entropy Inverse-Reinforcement Learning learns feature weights from demonstration, and like other maximum entropy models, offers strong performance guarantees and analysis possibilities. Proven generalization bounds are available that allow an estimate on the number of observed samples needed for a given expected level of error. However, they require knowledge of the covering number or complexity of the feature functions and/or known limits on the feature weights. When features are automatically extracted from a robot's sensor stream it is likely that many spurious features will be selected for use which could greatly increase the estimated number of samples needed, rendering the technique impractical. This project is developing an iterative, online variant of the maximum causal inverse-reinforcement learning algorithm that runs during the demonstrations and selects high-valued features as a critical subset which are then used to calculate the sample bounds. Once the required number of samples have been observed an offline inverse-reinforcement learning technique is run to ensure the feature weights are learned accurately. The new algorithm will be evaluated on a robot tasked with sorting previously-unknown objects. In this task, students will demonstrate the sorting of objects, then the robot will be required to do the same. Afterwards, the robot will be reset and the experiment repeats with a new set of objects. Critically, the software on the robot should not be changed or updated between these tasks.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
破骨细胞源性FcγRI介导类风湿性关节炎炎症后疼痛的作用机制
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    阳林
  • 依托单位:
四神丸调控生物钟基因Bmal1/Fc εRI介导肥大细胞节律性活化治疗IBS-D“晨起痛”的作用机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    何心凌
  • 依托单位:
NSUN6介导的m5C修饰调控心肌细胞凋亡和铁死亡参与MI/RI的机制研究
  • 批准号:
    2026JJ80739
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    袁乐宏
  • 依托单位:
中药牛耳枫中抗MI/RI新颖虎皮楠生物碱的发现与作用机制研究
  • 批准号:
    2026JJ60255
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    张济辉
  • 依托单位: