SMURF: SeMantic and linguistic UndeRstanding Fusion for Caption Evaluation via Typicality Analysis

SMURF: SeMantic and linguistic UndeRstanding Fusion for Caption Evaluation via Typicality Analysis
复制标题

DOI:
10.18653/v1/2021.acl-long.175
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Joshua Forster Feinglass;Yezhou Yang
Joshua Forster Feinglass;Yezhou Yang
中科院分区:
其他
文献类型:
--
作者:
Joshua Forster Feinglass;Yezhou Yang

文献摘要

被引文献

相似文献

视觉字幕的开放性使其成为一个具有挑战性的评价领域。大多数提出的模型依赖于专门的训练来提高人类的相关性,导致有限的采用,可推广性和可解释性。我们引入“典型性”,一个新的配方的评价植根于信息理论,这是唯一适合于缺乏一个明确的地面真相的问题。典型性作为我们的框架来开发一种新的语义比较,SPARCS,以及无参考流畅性评估指标。在我们的分析过程中,流利性的两个独立维度自然出现:风格,由度量SPURTS捕获,语法,以语法异常值惩罚的形式捕获。通过对基准数据集的广泛实验和消融研究,我们展示了这些语义和流畅性的分解维度如何为字幕差异提供更大的系统级洞察。与其他基于规则的评估指标相比,我们提出的指标及其组合SMURF沿着,实现了与人类判断的最新相关性。
The open-ended nature of visual captioning makes it a challenging area for evaluation. The majority of proposed models rely on specialized training to improve human-correlation, resulting in limited adoption, generalizability, and explainabilty. We introduce “typicality”, a new formulation of evaluation rooted in information theory, which is uniquely suited for problems lacking a definite ground truth. Typicality serves as our framework to develop a novel semantic comparison, SPARCS, as well as referenceless fluency evaluation metrics. Over the course of our analysis, two separate dimensions of fluency naturally emerge: style, captured by metric SPURTS, and grammar, captured in the form of grammatical outlier penalties. Through extensive experiments and ablation studies on benchmark datasets, we show how these decomposed dimensions of semantics and fluency provide greater system-level insight into captioner differences. Our proposed metrics along with their combination, SMURF, achieve state-of-the-art correlation with human judgment when compared with other rule-based evaluation metrics.