TBCMulti-Agent Reinforcement Learning for Assistive Robots
TBCMulti-Agent Reinforcement Learning for Assistive Robots
批准号:
2901369
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --
中文摘要
这个项目使用机器人控制的强化学习来帮助残疾人完成各种日常辅助任务。这些辅助性任务包括洗衣、穿衣、吃饭和喝水。该项目使用辅助健身环境,模拟这些任务,尽可能接近物理现实的设置(埃里克森等人,2019)。这个项目特别具有挑战性,因为它有很大的动作空间和很长的动作序列。此外,该项目强调人机交互,机器人需要学习能够满足人类偏好的策略,并预测人类的合作行为。目前,该项目正在复制和扩展现有研究中使用的基线算法。特别地,该项目将探索决策转换器在特定环境中的应用,这是一个仍然需要实现的体系结构实现。这是由短期记忆架构的成功所指导的,短期记忆架构更善于学习环境中的顺序依赖关系(Glaese et al., 2022)。变形金刚是LSTM的一个很有前途的替代品,它在学习长期依赖关系方面更有效,允许机器人在序列的早期选择更好的动作,以改善整个轨迹。未来的工作将侧重于挑战过去研究中普遍存在的限制性假设。特别是,该项目将解决两个假设:人类将与机器人进行最佳合作,人类的偏好是事先给定的,不会动态变化。为了消除这些假设,该项目将从人类反馈中实施强化学习,以适应更复杂和多样化的偏好集,而这些偏好集不需要预先定义。从人类数据中进行逆强化学习等技术可以模拟次优的人类合作。这些增强旨在开发算法,为实际应用程序学习更熟练的策略。
英文摘要
This project uses reinforcement learning for robotic control to aid disabled humans in various everyday assistive tasks. These assistive tasks include washing, dressing, eating and drinking. The project uses the assistive gym environment, which simulates these tasks to be as close as possible to physically realistic settings (Erickson et al., 2019). The project is particularly challenging as it has a large action space and long sequences of actions. Additionally, the project emphasises human-robot interaction, where the robot needs to learn policies that can satisfy the preferences of humans and anticipate the cooperative behaviour of humans. Currently, the project is replicating and expanding on baseline algorithms used in existing research. In particular, the project will explore the application of decision transformers to the specific environment, an architectural implementation that still needs to be implemented. This is guided by the success of short-term memory architectures that are better at learning sequential dependencies in the environment (Glaese et al., 2022). Transformers are a promising alternative to LSTM, being more efficient at learning long-term dependencies, allowing the robot to choose better action early on in the sequence to improve the entire trajectory. Future work will focus on challenging restrictive assumptions prevalent in past research. In particular, the project will address two assumptions: that humans will optimally cooperate with the robot and that human preferences are given ex-ante and do not dynamically change. To remove these assumptions, the project will implement reinforcement learning from human feedback, accommodating a more complex and diverse set of preferences that need not be predefined. Techniques like inverse reinforcement learning from human data can simulate sub-optimal human cooperation. These enhancements aim to develop algorithms that learn more adept policies for real-world applications.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
基于多模态 AI Agent的面部痤疮瘢痕临床特征评估与治疗方案优化系统的研究
-
批准号:2026JJ82357
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:罗滔
-
依托单位:
基于首创感染性疾病智能体UNION-Agent的SFTS全流程智慧管理模式探索性研究
-
批准号:JCZRQNB202600735
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:
-
依托单位:
基于Agent的自动化渗透测试技术研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:谭劲松
-
依托单位:
AI Agent赋能中小企业智能决策系统研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:蔡孝成
-
依托单位:
计算机控制Agent在可交互式企业征信报告生成的应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:林嘉诚
-
依托单位:
大模型Agent驱动的AI制药关键技术研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:
-
依托单位:
混合多元区域情境下多Agent的自主协同决策方法研究
-
批准号:62306099
-
项目类别:青年科学基金项目
-
资助金额:30.00万元
-
批准年份:2023
-
负责人:艾兵
-
依托单位:
基于操控员情境意识状态可解释Agent的智能交互触发机制研究
-
批准号:62376220
-
项目类别:面上项目
-
资助金额:50万元
-
批准年份:2023
-
负责人:于薇薇
-
依托单位:
基于多Agent仿真模型的新能源汽车市场渗透研究
-
批准号:2023JJ60196
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2023
-
负责人:黄建
-
依托单位:
面向联排联调的城市复合洪涝灾害风险Agent建模与智能决策
-
批准号:42371092
-
项目类别:面上项目
-
资助金额:50万元
-
批准年份:2023
-
负责人:王慧敏
-
依托单位: