Evaluating and Improving Factuality in Multimodal Abstractive Summarization

Evaluating and Improving Factuality in Multimodal Abstractive Summarization
复制标题

DOI:
10.48550/arxiv.2211.02580
复制
发表时间:
2022-11
期刊:
ArXiv
影响因子:
--
通讯作者:
David Wan;Mohit Bansal
David Wan;Mohit Bansal
中科院分区:
其他
文献类型:
--
作者:
David Wan;Mohit Bansal

文献摘要

被引文献

相似文献

目前用于评估抽象文档摘要的真实性的度量已经与人类判断实现了高度相关性,但是它们不考虑视觉模态,因此不足以用于视觉和语言摘要。我们提出了CLIPBERTSCORE,CLIPScore和BERTScore的简单加权组合,分别利用图像摘要和文档摘要之间的鲁棒性和强真实性检测性能。接下来,由于缺乏元评估基准来评估多模态真实性指标的质量,我们收集了人类对文档和图像真实性的判断。我们表明,这两个指标的简单组合在零拍设置实现了更高的相关性比现有的真实性指标的文档摘要,优于现有的多模态摘要指标,并执行竞争力强的多模态真实性指标专门微调的任务。我们的深入分析表明,CLIPBERTSCORE及其组件的鲁棒性和高相关性的四个真实性度量评估基准。最后,我们展示了我们的CLIPBERTSCORE度量的两个实际下游应用:用于选择在训练过程中关注的重要图像,以及作为强化学习的奖励,以提高多模态摘要生成的真实性。
Current metrics for evaluating factuality for abstractive document summarization have achieved high correlations with human judgment, but they do not account for the vision modality and thus are not adequate for vision-and-language summarization. We propose CLIPBERTSCORE, a simple weighted combination of CLIPScore and BERTScore to leverage the robustness and strong factuality detection performance between image-summary and document-summary, respectively. Next, due to the lack of meta-evaluation benchmarks to evaluate the quality of multimodal factuality metrics, we collect human judgments of factuality with respect to documents and images. We show that this simple combination of two metrics in the zero-shot setting achieves higher correlations than existing factuality metrics for document summarization, outperforms an existing multimodal summarization metric, and performs competitively with strong multimodal factuality metrics specifically fine-tuned for the task. Our thorough analysis demonstrates the robustness and high correlation of CLIPBERTSCORE and its components on four factuality metric-evaluation benchmarks. Finally, we demonstrate two practical downstream applications of our CLIPBERTSCORE metric: for selecting important images to focus on during training, and as a reward for reinforcement learning to improve factuality of multimodal summary generation w.r.t automatic and human evaluation.