课题基金 / 基金详情

TBCMulti-Agent Reinforcement Learning for Assistive Robots

TBCMulti-Agent Reinforcement Learning for Assistive Robots
TBC辅助机器人多智能体强化学习
批准号:
2901369
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
这个项目使用机器人控制的强化学习来帮助残疾人完成各种日常辅助任务。这些辅助性任务包括洗衣、穿衣、吃饭和喝水。该项目使用辅助健身环境,模拟这些任务,尽可能接近物理现实的设置(埃里克森等人,2019)。这个项目特别具有挑战性,因为它有很大的动作空间和很长的动作序列。此外,该项目强调人机交互,机器人需要学习能够满足人类偏好的策略,并预测人类的合作行为。目前,该项目正在复制和扩展现有研究中使用的基线算法。特别地,该项目将探索决策转换器在特定环境中的应用,这是一个仍然需要实现的体系结构实现。这是由短期记忆架构的成功所指导的,短期记忆架构更善于学习环境中的顺序依赖关系(Glaese et al., 2022)。变形金刚是LSTM的一个很有前途的替代品,它在学习长期依赖关系方面更有效,允许机器人在序列的早期选择更好的动作,以改善整个轨迹。未来的工作将侧重于挑战过去研究中普遍存在的限制性假设。特别是,该项目将解决两个假设:人类将与机器人进行最佳合作,人类的偏好是事先给定的,不会动态变化。为了消除这些假设,该项目将从人类反馈中实施强化学习,以适应更复杂和多样化的偏好集,而这些偏好集不需要预先定义。从人类数据中进行逆强化学习等技术可以模拟次优的人类合作。这些增强旨在开发算法,为实际应用程序学习更熟练的策略。
英文摘要
This project uses reinforcement learning for robotic control to aid disabled humans in various everyday assistive tasks. These assistive tasks include washing, dressing, eating and drinking. The project uses the assistive gym environment, which simulates these tasks to be as close as possible to physically realistic settings (Erickson et al., 2019). The project is particularly challenging as it has a large action space and long sequences of actions. Additionally, the project emphasises human-robot interaction, where the robot needs to learn policies that can satisfy the preferences of humans and anticipate the cooperative behaviour of humans. Currently, the project is replicating and expanding on baseline algorithms used in existing research. In particular, the project will explore the application of decision transformers to the specific environment, an architectural implementation that still needs to be implemented. This is guided by the success of short-term memory architectures that are better at learning sequential dependencies in the environment (Glaese et al., 2022). Transformers are a promising alternative to LSTM, being more efficient at learning long-term dependencies, allowing the robot to choose better action early on in the sequence to improve the entire trajectory. Future work will focus on challenging restrictive assumptions prevalent in past research. In particular, the project will address two assumptions: that humans will optimally cooperate with the robot and that human preferences are given ex-ante and do not dynamically change. To remove these assumptions, the project will implement reinforcement learning from human feedback, accommodating a more complex and diverse set of preferences that need not be predefined. Techniques like inverse reinforcement learning from human data can simulate sub-optimal human cooperation. These enhancements aim to develop algorithms that learn more adept policies for real-world applications.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
基于多模态 AI Agent的面部痤疮瘢痕临床特征评估与治疗方案优化系统的研究
基于首创感染性疾病智能体UNION-Agent的SFTS全流程智慧管理模式探索性研究
  • 批准号:
    JCZRQNB202600735
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
  • 依托单位:
基于Agent的自动化渗透测试技术研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    谭劲松
  • 依托单位:
AI Agent赋能中小企业智能决策系统研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    蔡孝成
  • 依托单位: