Storyboarding of Recipes: Grounded Contextual Generation

Storyboarding of Recipes: Grounded Contextual Generation
复制标题

食谱的故事板:扎根的情境生成

DOI:
10.18653/v1/p19-1606
复制
发表时间:
2019
期刊:
--
影响因子:
--
通讯作者:
A. Black
A. Black
中科院分区:
--
文献类型:
--
作者:
Khyathi Raghavi Chandu;Eric Nyberg;A. Black

文献摘要

参考文献

被引文献

相似文献

人类的信息需求本质上是多模态的,能够最大限度地利用情境。我们介绍了一个数据集的顺序程序(如何)文本生成图像在烹饪领域。该数据集由16,441个烹饪食谱组成,其中160,479张照片与不同步骤相关。我们设置了一个基线,其动机是在视觉故事讲述(ViST)任务的人类评估方面表现最好的模型。此外,我们引入了两种模型,将有限状态机(FSM)学习的高级结构纳入神经序列生成过程:(1)解码器中的支架结构(SSiD)(2)丢失中的支架结构(SSiL)。我们的最佳性能模型(SSiL)的METEOR得分为0.31,比基线模型提高了0.6。我们还对生成的基础食谱进行了人工评估,结果显示,61%的人认为我们提出的(SSiL)模型在整体食谱方面优于基线模型。我们还讨论了输出的分析,突出了未来方向的关键重要NLP问题。
Information need of humans is essentially multimodal in nature, enabling maximum exploitation of situated context. We introduce a dataset for sequential procedural (how-to) text generation from images in cooking domain. The dataset consists of 16,441 cooking recipes with 160,479 photos associated with different steps. We setup a baseline motivated by the best performing model in terms of human evaluation for the Visual Story Telling (ViST) task. In addition, we introduce two models to incorporate high level structure learnt by a Finite State Machine (FSM) in neural sequential generation process by: (1) Scaffolding Structure in Decoder (SSiD) (2) Scaffolding Structure in Loss (SSiL). Our best performing model (SSiL) achieves a METEOR score of 0.31, which is an improvement of 0.6 over the baseline model. We also conducted human evaluation of the generated grounded recipes, which reveal that 61% found that our proposed (SSiL) model is better than the baseline model in terms of overall recipes. We also discuss analysis of the output highlighting key important NLP issues for prospective directions.
DOI: 10.1609/aaai.v33i01.33016949
发表时间: 2018-11
期刊: PLoS ONE
影响因子: 3.7
作者:
Victor Sanh;Thomas Wolf;Sebastian Ruder
通讯作者: Victor Sanh;Thomas Wolf;Sebastian Ruder