Physics-Informed Reinforcement Learning for Tactile Manipulation
Physics-Informed Reinforcement Learning for Tactile Manipulation
批准号:
2614951
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
在非结构化和不熟悉的环境中进行对象操作是一种强大的工具,可以大大提高机器人技术的价值。然而,在这一领域的研究工作主要集中在使用视觉作为输入控制,而很少注意到触摸的应用。触觉传感是实现人类水平灵活性的关键。即使在物体被部分遮挡时,也可以获得接触顺应性、接触位置、剪切力、压力、滑动和接触的其他重要机械特性等测量结果,否则这些测量结果将无法通过视觉获得。在我的研究中,我希望利用这种触摸的力量来开发控制框架,以支持机器人与现实世界的互动。为了充分实现自主操作的价值,机器人需要在广泛的普遍和非系统的场景下操作,这是它和它的设计者以前都没有预见到的,其中许多是不可能建模的。这就需要具有高度通用性和通用性的高级控制框架,这些框架也足够强大,可以处理一系列不可预测的情况。强化学习(RL)是一种流行的学习算法,可以让机器人通过试错来学习复杂的控制器。采用这种技术将简化控制器设计问题,避免为我们非结构化世界中的每一种场景开发专门的控制器,并显着减少实现人类水平灵活性所需的研究工作。然而,强化学习可能非常脆弱。学习过程可能对某些设计参数非常敏感,与预期行为的微小偏差可能导致灾难性故障。它还具有高样本复杂性,这使得现实世界的学习不可行。这些缺点意味着机器人强化学习的许多最新进展都涉及在转移到现实世界之前进行模拟实验,从而产生许多模拟到真实的问题。基于模型的强化学习是一种较新的方法,旨在通过首先学习过渡动态的表示来解决这些缺点,然后使用它来导出最优控制器。这已被证明是更有效的机器人控制,其中高容量的模型,如神经网络或概率模型可以用来解决复杂的控制问题,减少相互作用的数量。这不仅使现实世界的学习成为可能,而且模型的存在可以提高鲁棒性以及许多其他好处。基于模型的强化学习在触觉机器人中的应用还没有得到充分的探索,我相信这种结合可以大大提高触觉操纵的能力。我的研究将致力于进一步实现这一愿景,重点是安全有效的现实世界学习和部署。我希望开发用于合成控制器的方法,即:在在线学习过程中有效地采样以表现出自适应行为,在部署中具有鲁棒性以能够处理不可预测的情况,并且在制定中也保持通用性,以便它可以应用于一系列触觉操作任务。我希望探索的一个潜在方向是结合物理先验来指导学习过程,这可能会减轻与强化学习的试错性质相关的主要效率问题。这项研究属于EPSRC人工智能和机器人领域,福尔斯。
英文摘要
Object manipulation in unstructured and unfamiliar environments is a powerful tool that can greatly increase the value of robotics. However, research effort in this area has been largely focused on using vision as input for control whilst little attention has been given to the application of touch. Tactile sensing is key for achieving human-level dexterity. Measurements such as contact compliance, contact location, shear, pressure, slip, and other important mechanical properties of contact, that would otherwise be inaccessible through vision, can be obtained even when objects are partially occluded. In my research, I hope to harness this power of touch to develop control frameworks to support robots' interaction with the real-world. To fully realise the value of autonomous manipulation, robots need to operate under a wide range of pervasive and unsystematic scenarios that it nor its designers have foreseen before, many of which can be impossible to model. This calls for advanced control frameworks with high versatility and generality that is also robust enough to deal with a range of unpredictable situations. Reinforcement learning (RL) is a popular learning algorithm that can allow robots to learn a complex controller through trial and error. Adopting such technique would simplify the controller design problem, avoiding the need to develop specialised controller for every scenario in our unstructured world and significantly reducing the research effort required to achieve human-level dexterity. However, reinforcement learning can be extremely fragile. The learning process can be very sensitive to certain design parameters and small deviations from expected behaviour can lead to catastrophic failures. It also suffers from high sample complexity which makes real-world learning unfeasible. These drawbacks have meant that much of the recent advances in reinforcement learning for robotics have involved experimenting in simulations before transferring to the real-world, creating many sim-to-real problems. Model-based reinforcement learning is a more recent approach that aims to tackle these shortcomings by first learning a representation of the transition dynamics before using it to derive the optimal controller. This has been shown to be more effective for robot control where high-capacity models such as neural networks or probabilistic models can be used to solve complex control problems with reduced number of interactions. This not only makes real-world learning possible, but the existence of a model can improve robustness amongst many other benefits. The application of model-based reinforcement learning to tactile robotics has very much been underexplored and I believe this combination could massively progress the capabilities of tactile manipulation. My research will aim to further this vision, with a focus on safe and efficient real-world learning and deployment. I hope to develop methodologies for synthesising controller that is; sample efficient in the online learning process to exhibit adaptive behaviour, robust in deployment to be able to deal with unpredictable situations, and also remain general in formulation such that it can be applied to a range of tactile manipulation tasks. A potential direction I hope to explore is to incorporate physics-priors to guide the learning process which can potentially mitigate the major efficiency issues associated with the trial-and-error nature of reinforcement learning. This research falls within the EPSRC artificial intelligence and robotics area.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金