SGL: Symbolic Goal Learning in a Hybrid, Modular Framework for Human Instruction Following

SGL: Symbolic Goal Learning in a Hybrid, Modular Framework for Human Instruction Following
复制标题

DOI:
10.1109/lra.2022.3190076
复制
发表时间:
2022-02
影响因子:
5.2
通讯作者:
Ruinian Xu;Hongyi Chen;Yunzhi Lin;P. Vela
Ruinian Xu;Hongyi Chen;Yunzhi Lin;P. Vela
中科院分区:
计算机科学2区
文献类型:
--
作者:
Ruinian Xu;Hongyi Chen;Yunzhi Lin;P. Vela

文献摘要

相似文献

本文探讨了人类的指令通过一个混合的,模块化的系统与符号和连接元素的机器人操作。符号方法构建具有语义解析和任务规划模块的模块化系统,用于从自然语言请求中产生动作序列。现代连接主义方法采用深度神经网络,学习视觉和语言特征,以端到端的方式将输入映射到一系列低级别动作。混合模块化系统融合了这两种方法,创建了一个模块化框架:它通过深度神经网络将指令制定为符号目标学习,然后通过符号规划器进行任务规划。连接主义和符号模块与规划领域定义语言连接在一起。视觉和语言学习网络预测其目标表示,并将其发送到规划器以产生完成任务的动作序列。为了提高自然语言的灵活性,我们进一步将隐含的人类意图与显式的人类指令结合起来。为了学习视觉和语言的通用特征,我们建议在场景图解析和语义文本相似性任务上分别预训练视觉和语言编码器。基准评估的影响,不同的组件,或选项,视觉和语言学习模式,并显示预培训策略的有效性。在模拟器AI2THOR中进行的操作实验表明,该框架对新的场景具有鲁棒性。
This paper investigates human instruction following for robotic manipulation via a hybrid, modular system with symbolic and connectionist elements. Symbolic methods build modular systems with semantic parsing and task planning modules for producing sequences of actions from natural language requests. Modern connectionist methods employ deep neural networks that learn visual and linguistic features for mapping inputs to a sequence of low-level actions, in an end-to-end fashion. The hybrid, modular system blends these two approaches to create a modular framework: it formulates instruction following as symbolic goal learning via deep neural networks followed by task planning via symbolic planners. Connectionist and symbolic modules are bridged with Planning Domain Definition Language. The vision-and-language learning network predicts its goal representation, which is sent to a planner for producing a task-completing action sequence. For improving the flexibility of natural language, we further incorporate implicit human intents with explicit human instructions. To learn generic features for vision and language, we propose to separately pretrain vision and language encoders on scene graph parsing and semantic textual similarity tasks. Benchmarking evaluates the impacts of different components of, or options for, the vision-and-language learning model and shows the effectiveness of pretraining strategies. Manipulation experiments conducted in the simulator AI2THOR show the robustness of the framework to novel scenarios.