State-aware video procedural captioning

State-aware video procedural captioning
复制标题

DOI:
10.1007/s11042-023-14774-7
复制
发表时间:
2021-10
影响因子:
3.6
通讯作者:
Taichi Nishimura;Atsushi Hashimoto;Y. Ushiku;Hirotaka Kameko;Shinsuke Mori
Taichi Nishimura;Atsushi Hashimoto;Y. Ushiku;Hirotaka Kameko;Shinsuke Mori
中科院分区:
计算机科学4区
文献类型:
--
作者:
Taichi Nishimura;Atsushi Hashimoto;Y. Ushiku;Hirotaka Kameko;Shinsuke Mori

文献摘要

相似文献

视频过程字幕(VPC)是从教学视频中生成过程文本的一种方法,是场景理解和实际应用的一项重要任务。VPC的主要挑战是描述如何准确地操纵材料。本文通过设计一个新的VPC任务,从教学视频和材料集的剪辑序列中生成程序文本来关注这一挑战。在该任务中,材料的状态通过操纵顺序地改变,产生它们的状态感知视觉表示(例如,鸡蛋被转变成破裂、搅拌、然后油炸的形式)。最大的困难是将这种视觉表示转换为文本表示;也就是说,模型应该在操作后跟踪材料状态,以更好地关联跨模态关系。为了实现这一目标,我们提出了一种新的VPC方法,它修改了现有的文本模拟器跟踪材料状态作为视觉模拟器,并将其纳入视频字幕模型。我们的实验结果表明了该方法的有效性,其性能优于最先进的视频字幕模型。我们进一步分析了学习嵌入的材料,以证明模拟器捕捉他们的状态转换。
Video procedural captioning (VPC), which generates procedural text from instructional videos, is an essential task for scene understanding and real-world applications. The main challenge of VPC is to describe how to manipulate materials accurately. This paper focuses on this challenge by designing a new VPC task, generating a procedural text from the clip sequence of an instructional video and material set. In this task, the state of materials is sequentially changed by manipulations, yielding their state-aware visual representations (e.g., eggs are transformed into cracked, stirred, then fried forms). The essential difficulty is to convert such visual representations into textual representations; that is, a model should track the material states after manipulations to better associate the cross-modal relations. To achieve this, we propose a novel VPC method, which modifies an existing textual simulator for tracking material states as a visual simulator and incorporates it into a video captioning model. Our experimental results show the effectiveness of the proposed method, which outperforms state-of-the-art video captioning models. We further analyze the learned embedding of materials to demonstrate that the simulators capture their state transition.