课题基金 / 基金详情

Sim-to-Real Deep Reinforcement Learning for legged robot locomotion with vision-based high dimensional data

Sim-to-Real Deep Reinforcement Learning for legged robot locomotion with vision-based high dimensional data
使用基于视觉的高维数据进行腿式机器人运动的模拟到真实深度强化学习
批准号:
1950742
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
本博士课程将探索让有腿机器人改进和适应不同地形的方法。对于MSc,物理模拟器预先编程了环境地形和机器人尺寸。对摩擦系数和重量分布等参数进行了粗略估计。然而,对于博士学位,机器人将在3D环境中建立自己的模型,使用视觉和深度感知结合方向感知和机器人牙牙学语。这将使机器人能够调整模拟环境的参数,从而增加适应性,减少与现实的差距。目的是提供一种新颖的方法,允许任何类型的有腿机器人从高维输入数据中进行操作。该方法旨在解决显式编程算法的适应性问题,同时也解决了PPO强化学习的现实差距问题。输入数据将来自RGBD、关节参数、方向和触觉感应。注意,PPO在处理高维输入时表现得非常好,为了识别现实世界的复杂性,这将是必需的。宗旨和目标:1 .构建一个能够感知环境的有腿机器人(即RGBD,方向和触觉传感器)自我建模代理-允许机器人进行机器人牙牙学语,并使用方向、触觉和关节位置传感器对代理进行自我建模。模拟环境——允许机器人用RGBD传感器扫描房间,模拟其地形并识别其目标位置(例如房间中的球)在三维世界中用强化学习训练模拟机器人,以确定到达目标位置的策略5。将训练好的策略部署到物理机器人上。测量机器人在模拟和现实世界中的表现之间的“现实差距”。根据机器人的咿呀学语和环境建模进行调整。TRL 4四足机器人,可以穿越以前看不见的地形/场景到达目标
英文摘要
This PhD will explore methods that allow legged robots to improve and adapt its gaits to various terrains. For the MSc the physics simulator was pre-programmed with the environment terrain and robot dimensions. Parameters such as friction coefficients and weight distributions were roughly estimated. However, for the PhD the robot will build a model of itself in a 3D environment using a combination of vision and depth sensing combined with orientation sensing and robot babbling. This will allow the robot to adjust the parameters of the simulated environment which may increase adaptability and reduce the reality gap.The aim is to contribute a novel method that allows any type of legged robot to manoeuvre from high dimensional input data. This method aims to address the adaptability problems with explicitly programmed algorithms whist also addressing the reality gap issues with PPO reinforcement learning. The input data will be from RGBD, joint parameters, orientation and tactile sensing. Note PPO performs exceptionally well with high dimension inputs and these will be required in order to identify the complexities of the real world.Aims and objectives:1. Build a legged robot capable of sensing its environment (i.e. RGBD, orientation and tactile sensors)2. Self-model Agent - Allow the robot to perform robot babbling and use the orientation, tactile and joint position sensors to self-model the agent3. Model Environment - Allow the robot to scan the room with RGBD sensors to model its terrain and identify its target location (i.e. a ball in a room)4. Train the simulated robot with reinforcement learning in the 3D world to determine a policy to reach the target location5. Deploy the trained policy on the physical robot 6. Measure the 'reality gap' between the robot's performance in simulation and the physical world.7. Adapt robot babbling and environment modelling accordingly.8. TRL 4 quadruped robot that can traverse previously unseen terrains/scenarios toward a goal
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Immuno-Real Time PCR法精确定量血清MG7抗原及在早期胃癌预警中的价值
无色ReAl3(BO3)4(Re=Y,Lu)系列晶体紫外倍频性能与器件研究