State-aware video procedural captioning
State-aware video procedural captioning
复制标题
DOI:
10.1007/s11042-023-14774-7
复制
发表时间:
2021-10
影响因子:
3.6
通讯作者:
Taichi Nishimura;Atsushi Hashimoto;Y. Ushiku;Hirotaka Kameko;Shinsuke Mori
中科院分区:
文献类型:
--
作者:
Taichi Nishimura;Atsushi Hashimoto;Y. Ushiku;Hirotaka Kameko;Shinsuke Mori
Video procedural captioning (VPC), which generates procedural text from instructional videos, is an essential task for scene understanding and real-world applications. The main challenge of VPC is to describe how to manipulate materials accurately. This paper focuses on this challenge by designing a new VPC task, generating a procedural text from the clip sequence of an instructional video and material set. In this task, the state of materials is sequentially changed by manipulations, yielding their state-aware visual representations (e.g., eggs are transformed into cracked, stirred, then fried forms). The essential difficulty is to convert such visual representations into textual representations; that is, a model should track the material states after manipulations to better associate the cross-modal relations. To achieve this, we propose a novel VPC method, which modifies an existing textual simulator for tracking material states as a visual simulator and incorporates it into a video captioning model. Our experimental results show the effectiveness of the proposed method, which outperforms state-of-the-art video captioning models. We further analyze the learned embedding of materials to demonstrate that the simulators capture their state transition.