ImaginE: An Imagination-Based Automatic Evaluation Metric for Natural Language Generation

ImaginE: An Imagination-Based Automatic Evaluation Metric for Natural Language Generation
复制标题

DOI:
10.18653/v1/2023.findings-eacl.6
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Wanrong Zhu;X. Wang;An Yan;M. Eckstein;W. Wang
Wanrong Zhu;X. Wang;An Yan;M. Eckstein;W. Wang
中科院分区:
其他
文献类型:
--
作者:
Wanrong Zhu;X. Wang;An Yan;M. Eckstein;W. Wang

文献摘要

被引文献

相似文献

自然语言生成(NLG)的自动评估通常依赖于与文本引用的标记级或嵌入级比较。这与人类的语言处理不同,视觉想象通常会提高理解力。在这项工作中,我们提出了ImaginE,自然语言生成的一个基于解释的自动评价指标。借助StableDiffusion,一个最先进的文本到图像生成器,我们自动生成一个图像作为文本片段的体现想象力,并使用上下文嵌入计算想象力相似度。跨几个文本生成任务的实验表明,添加机器生成的图像与我们的ImaginE显示了巨大的潜力,在引入多模态信息到NLG评估,并提高现有的自动度量的相关性与人类相似性判断在基于参考和无参考的评估方案。
Automatic evaluations for natural language generation (NLG) conventionally rely on token-level or embedding-level comparisons with text references. This differs from human language processing, for which visual imagination often improves comprehension. In this work, we propose ImaginE, an imagination-based automatic evaluation metric for natural language generation. With the help of StableDiffusion, a state-of-the-art text-to-image generator, we automatically generate an image as the embodied imagination for the text snippet and compute the imagination similarity using contextual embeddings. Experiments spanning several text generation tasks demonstrate that adding machine-generated images with our ImaginE displays great potential in introducing multi-modal information into NLG evaluation, and improves existing automatic metrics’ correlations with human similarity judgments in both reference-based and reference-free evaluation scenarios.