Recipe Generation from Unsegmented Cooking Videos

Recipe Generation from Unsegmented Cooking Videos
复制标题

DOI:
10.48550/arxiv.2209.10134
复制
发表时间:
2022-09
期刊:
ACM Transactions on Multimedia Computing, Communications and Applications
影响因子:
--
通讯作者:
Taichi Nishimura;Atsushi Hashimoto;Y. Ushiku;Hirotaka Kameko;Shinsuke Mori
Taichi Nishimura;Atsushi Hashimoto;Y. Ushiku;Hirotaka Kameko;Shinsuke Mori
中科院分区:
其他
文献类型:
--
作者:
Taichi Nishimura;Atsushi Hashimoto;Y. Ushiku;Hirotaka Kameko;Shinsuke Mori

文献摘要

相似文献

本文解决了从未分段的烹饪视频中生成菜谱的问题,该任务要求智能体(1)提取完成菜肴的关键事件,以及(2)为提取的事件生成句子。我们的任务类似于密集视频字幕(DVC),旨在彻底检测事件并为其生成句子。然而,与 DVC 不同的是,在菜谱生成中,菜谱故事意识至关重要,模型应该以正确的顺序提取适当数量的事件,并根据它们生成准确的句子。我们分析了 DVC 模型的输出,并确认虽然 (1) 有几个事件可以作为食谱故事,(2) 为这些事件生成的句子并不以视觉内容为基础。基于此,我们的目标是通过从输出事件中选择预言事件并为其重新生成句子来获得正确的食谱。为了实现这一目标,我们提出了一种基于 Transformer 的多模态循环方法,用于训练事件选择器和句子生成器,以便从 DVC 的事件中选择预言事件并为其生成句子。此外,我们还通过包含成分来扩展模型,以生成更准确的食谱。实验结果表明,所提出的方法优于最先进的 DVC 模型。我们还确认,通过以故事感知的方式对菜谱进行建模,所提出的模型以正确的顺序输出适当数量的事件。
This paper tackles recipe generation from unsegmented cooking videos, a task that requires agents to (1) extract key events in completing the dish and (2) generate sentences for the extracted events. Our task is similar to dense video captioning (DVC), which aims at detecting events thoroughly and generating sentences for them. However, unlike DVC, in recipe generation, recipe story awareness is crucial, and a model should extract an appropriate number of events in the correct order and generate accurate sentences based on them. We analyze the output of the DVC model and confirm that although (1) several events are adoptable as a recipe story, (2) the generated sentences for such events are not grounded in the visual content. Based on this, we set our goal to obtain correct recipes by selecting oracle events from the output events and re-generating sentences for them. To achieve this, we propose a transformer-based multimodal recurrent approach of training an event selector and sentence generator for selecting oracle events from the DVC’s events and generating sentences for them. In addition, we extend the model by including ingredients to generate more accurate recipes. The experimental results show that the proposed method outperforms state-of-the-art DVC models. We also confirm that, by modeling the recipe in a story-aware manner, the proposed model outputs the appropriate number of events in the correct order.