What Should the System Do Next?: Operative Action Captioning for Estimating System Actions

What Should the System Do Next?: Operative Action Captioning for Estimating System Actions
复制标题

DOI:
10.48550/arxiv.2210.02735
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Taiki Nakamura;Seiya Kawano;Akishige Yuguchi;Yasutomo Kawanishi;Koichiro Yoshino
Taiki Nakamura;Seiya Kawano;Akishige Yuguchi;Yasutomo Kawanishi;Koichiro Yoshino
中科院分区:
其他
文献类型:
--
作者:
Taiki Nakamura;Seiya Kawano;Akishige Yuguchi;Yasutomo Kawanishi;Koichiro Yoshino

文献摘要

相似文献

- 机器人需要根据观察者的身份正确理解周围情况的人类辅助系统,并输出人类所需的支持动作是与人类交流的重要渠道之一。为了表达他们在这项研究中的理解和行动计划。我们构建了一个系统,该系统对可能的操作动作进行了口头描述,该操作动作将当前状态更改为给定的目标状态。标题是在日常生活情况下通过众包将当前状态变为目标状态的行动动作的标题有望包含一些改变状态的动作,我们将场景图预测用作辅助任务,因为在场景图中所写的事件与状态变化相对应。在当前和目标状态之间进行了预测的辅助任务。
— Such human-assisting systems as robots need to correctly understand the surrounding situation based on obser- vations and output the required support actions for humans. Language is one of the important channels to communicate with humans, and the robots are required to have the ability to express their understanding and action planning results. In this study, we propose a new task of operative action captioning that estimates and verbalizes the actions to be taken by the system in a human-assisting domain. We constructed a system that outputs a verbal description of a possible operative action that changes the current state to the given target state. We collected a dataset consisting of two images as observations, which express the current state and the state changed by actions, and a caption that describes the actions that change the current state to the target state, by crowdsourcing in daily life situations. Then we constructed a system that estimates operative action by a caption. Since the operative action’s caption is expected to contain some state-changing actions, we use scene-graph prediction as an auxiliary task because the events written in the scene graphs correspond to the state changes. Experimental results showed that our system successfully described the operative actions that should be conducted between the current and target states. The auxiliary tasks that predict the scene graphs improved the quality of the estimation results.