QAFactEval: Improved QA-Based Factual Consistency Evaluation for Summarization

QAFactEval: Improved QA-Based Factual Consistency Evaluation for Summarization
复制标题

QAFactEval:改进的基于 QA 的事实一致性评估的摘要

DOI:
10.18653/v1/2022.naacl-main.187
复制
发表时间:
2021
期刊:
ArXiv
影响因子:
--
通讯作者:
Caiming Xiong
Caiming Xiong
中科院分区:
--
文献类型:
--
作者:
Alexander R. Fabbri;C. Wu;Wenhao Liu;Caiming Xiong

文献摘要

参考文献

被引文献

相似文献

事实一致性是文本摘要模型在实际应用中的一个重要特性。现有的工作在评估这方面可以大致分为两条研究路线,蕴涵为基础的和问答(QA)为基础的指标,不同的实验设置往往会导致对比的结论,哪种范式表现最好。在这项工作中,我们进行了广泛的比较蕴涵和QA为基础的指标,表明仔细选择的组件的QA为基础的指标,特别是问题的生成和可回答性分类,是至关重要的性能。基于这些见解,我们提出了一个优化的指标,我们称之为QAFactEval,它比SummaC事实一致性基准上以前基于QA的指标平均提高了14%,并且优于性能最好的基于蕴涵的指标。此外,我们发现,基于QA和基于蕴涵的指标可以提供互补的信号,并结合成一个单一的指标,以进一步提高性能。
Factual consistency is an essential quality of text summarization models in practical settings. Existing work in evaluating this dimension can be broadly categorized into two lines of research, entailment-based and question answering (QA)-based metrics, and different experimental setups often lead to contrasting conclusions as to which paradigm performs the best. In this work, we conduct an extensive comparison of entailment and QA-based metrics, demonstrating that carefully choosing the components of a QA-based metric, especially question generation and answerability classification, is critical to performance. Building on those insights, we propose an optimized metric, which we call QAFactEval, that leads to a 14% average improvement over previous QA-based metrics on the SummaC factual consistency benchmark, and also outperforms the best-performing entailment-based metric. Moreover, we find that QA-based and entailment-based metrics can offer complementary signals and be combined into a single metric for a further performance boost.
DOI: 10.18653/v1/p19-1213
发表时间: 2019-05
期刊: --
影响因子: --
作者:
Tobias Falke;Leonardo F. R. Ribeiro;Prasetya Ajie Utama;Ido Dagan;Iryna Gurevych
通讯作者: Tobias Falke;Leonardo F. R. Ribeiro;Prasetya Ajie Utama;Ido Dagan;Iryna Gurevych