DocAsRef: An Empirical Study on Repurposing Reference-based Summary Quality Metrics as Reference-free Metrics

DocAsRef: An Empirical Study on Repurposing Reference-based Summary Quality Metrics as Reference-free Metrics
复制标题

DOI:
10.18653/v1/2023.findings-emnlp.87
复制
发表时间:
2022-12
期刊:
--
影响因子:
--
通讯作者:
F. Bao;Ruixuan Tu;Ge Luo
F. Bao;Ruixuan Tu;Ge Luo
中科院分区:
其他
文献类型:
--
作者:
F. Bao;Ruixuan Tu;Ge Luo

文献摘要

相似文献

自动摘要质量评估福尔斯分为两类:基于参考的和无参考的。基于参考的指标,历史上被认为是更准确的,由于人类书面参考提供的额外信息,受到其对人类输入的依赖的限制。在本文中,我们假设,比较方法所使用的一些基于参考的指标,以评估其相应的参考系统摘要,可以有效地适应评估它对它的源文件,从而将这些指标转化为无参考的。实验结果支持这一假设。在重新使用参考后,使用<0.5B参数的预训练DeBERTa-large-MNLI模型的零杆BERTScore在SummEval和Newsroom数据集上的各个方面始终优于其原始的基于参考的版本。与大多数现有的无参考指标相比,它也表现出色,并与基于GPT-3.5的零射击摘要评估器展开激烈竞争。
Automated summary quality assessment falls into two categories: reference-based and reference-free. Reference-based metrics, historically deemed more accurate due to the additional information provided by human-written references, are limited by their reliance on human input. In this paper, we hypothesize that the comparison methodologies used by some reference-based metrics to evaluate a system summary against its corresponding reference can be effectively adapted to assess it against its source document, thereby transforming these metrics into reference-free ones. Experimental results support this hypothesis. After being repurposed reference-freely, the zero-shot BERTScore using the pretrained DeBERTa-large-MNLI model of<0.5B parameters consistently outperforms its original reference-based version across various aspects on the SummEval and Newsroom datasets. It also excels in comparison to most existing reference-free metrics and closely competes with zero-shot summary evaluators based on GPT-3.5.