Efficient Robotic Reinforcement Learning via Off-Policy and Meta-Learning
Efficient Robotic Reinforcement Learning via Off-Policy and Meta-Learning
批准号:
2285275
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --
中文摘要
该项目属于EPSRC人工智能和机器人研究领域的福尔斯。研究ContextDeep基于强化学习的方法越来越多地被研究用于机器人技术,因为它们承诺提供更灵活的控制策略,减少人工工程开销。与传统的机器人方法不同,在传统的机器人方法中,控制策略是由高度专业化的专家分别为每个任务指定的,学习算法可以从自己的经验中获得一般行为,就像许多生物有机体一样。深度学习模型在不同的数据集上训练时表现出良好的泛化能力,其成功的关键在于它们能够从大量训练数据中学习数百万个参数。现实世界中机器人学习的一个主要限制是,我们无法在单个实验中收集足够大的数据集来进行“ImageNet规模”的泛化。潜在影响虽然普通消费者越来越负担得起机器人,但由于难以设计鲁棒的控制策略,它们可以执行的任务集有限。机器人可以减轻人类在许多日常任务中的负担,例如在家做饭,在辅助生活社区照顾老人,在医院做手术,或在危险的灾区进行救援行动。能够有效学习和推广的算法对于向更广泛的受众传播有用的低成本机器人至关重要。目标和研究方法对于强化学习算法演变为复杂现实任务的实用方法,我们必须设计新颖的算法,使我们能够绕过数据稀缺的问题。一种可能的方法是更好地利用现有的历史数据。为了实现这一目标,我们建议研究如何(1)更好地利用政策外数据,即在特定机器人实验之外收集的数据,以及(2)可以快速适应新任务的元学习策略。有大量以前收集的机器人数据可用,这已经为学习机器人方法提供了大量多样的经验。有了将这种经验融入强化学习的能力,我们可以让策略真正在不同的对象、环境、场景甚至不同的机器人之间泛化。例如,我们可以使用RoboNet数据集来改进单任务或多任务强化学习的训练。我们将手动定义一个或多个任务,并使用这些任务奖励重新标记所有RoboNet数据。然后,我们将在RoboNet的大数据集和每个任务的适量新数据上运行一个非策略强化算法,例如Soft-Actor Critic。元强化学习算法允许代理通过利用先前收集的经验的结构相似性来快速适应新任务。现有的元学习算法主要是在所有经验都可以在一个单一批次中被学习者访问的环境中操作。更现实的是,在真实的世界中,代理遇到的任务通常是以顺序的方式体验的,这就是为什么我们应该扩展当前的元学习公式来支持这种流体验的情况。另一个有趣的方向是制定Meta强化学习的版本,其中所有的MDP不一定共享相同的状态和动作空间。这需要开发新的模型架构,可以读取异构状态空间并输出异构操作。
英文摘要
This project falls within the EPSRC Artificial Intelligence and Robotics research areas.Research ContextDeep reinforcement learning-based methods are increasingly researched approaches for robotics because of their promise to provide more flexible control policies with reduced manual engineering overhead. In contrast to traditional robotics methods, where the control policies are specified by highly specialized experts for each task separately, learning algorithms can acquire general behaviors from their own experience in the same way that many biological organisms do. Deep learning models are shown to generalize well when trained on diverse datasets, and the key to their success lies in their ability to learn millions of parameters from large amounts of training data. One of the major limitations of real-world robotic learning is that we cannot afford to collect large enough datasets for "ImageNet-scale" generalization within a single experiment.Potential ImpactWhile robots are becoming increasingly affordable to average consumers, the set of tasks they can carry out is limited due to the difficulty of designing robust control policies. Robots could reduce the human burden in many everyday tasks such as cooking at homes, elderly care at assisted-living communities, surgery at hospitals, or rescue operations in dangerous disaster zones. Algorithms that can learn and generalize efficiently are crucial to disseminating useful low-cost robots for wider audiences.Objectives and Research MethodologyFor reinforcement learning algorithms to evolve into practical methods for complex real-world tasks, we must design novel algorithms that allow us to get around the issue of data scarcity. One possible way is to better leverage existing historical data. Towards this goal, we propose to investigate how to (1) better utilize off-policy data, that is, the data collected outside of the specific robot experiment, and (2) meta-learn policies that can adapt to new tasks quickly.There is an abundance of previously collected robotic data available, which already provides a large and diverse experience for learning robotic methods. With the ability to incorporate this experience into reinforcement learning, we can get the policies to truly generalize across different objects, environments, scenes, and possibly even across different robots. For example, we could use the RoboNet dataset to improve the training of a single-task or multi-task reinforcement learning. We would define one or more tasks manually and relabel all the RoboNet data with these task rewards. We would then run an off-policy reinforcement algorithm, such as Soft-Actor Critic, on the large set of RoboNet data and a modest amount of new data for each task. This should allow policies to generalize and learn faster.Meta-reinforcement learning algorithms allow agents to rapidly adapt to new tasks by exploiting the structural similarities of previously collected experiences. Existing meta-learning algorithms operate mainly in a setting where all the experience is accessible to the learner in a single batch. More realistically, in the real world, the tasks encountered by agents are typically experienced in a sequential fashion, which is why we should extend the current meta-learning formulation to support such cases of streaming experiences. Another interesting direction would be to formulate versions of meta reinforcement learning where all the MDPs don't necessarily share the same state and action spaces. This would require developing new model architectures that can read in heterogeneous state spaces and output heterogeneous actions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
High-precision force-reflected bilateral teleoperation of multi-DOF hydraulic robotic manipulators
-
批准号:52111530069
-
项目类别:国际(地区)合作与交流项目
-
资助金额:10万元
-
批准年份:2021
-
负责人:徐兵
-
依托单位: