Ranking Generated Summaries by Correctness: An Interesting but Challenging Application for Natural Language Inference

Ranking Generated Summaries by Correctness: An Interesting but Challenging Application for Natural Language Inference
复制标题

DOI:
10.18653/v1/p19-1213
复制
发表时间:
2019-05
期刊:
--
影响因子:
--
通讯作者:
Tobias Falke;Leonardo F. R. Ribeiro;Prasetya Ajie Utama;Ido Dagan;Iryna Gurevych
Tobias Falke;Leonardo F. R. Ribeiro;Prasetya Ajie Utama;Ido Dagan;Iryna Gurevych
中科院分区:
其他
文献类型:
--
作者:
Tobias Falke;Leonardo F. R. Ribeiro;Prasetya Ajie Utama;Ido Dagan;Iryna Gurevych

文献摘要

被引文献

相似文献

虽然最近的进展,抽象的摘要已经导致非常流畅的摘要,在生成的摘要中的事实错误仍然严重限制了其在实践中的使用。在本文中,我们评估了通过众包由最先进的模型产生的摘要,并表明这种错误经常发生,特别是在更抽象的模型中。我们研究是否可以使用文本蕴涵预测来检测这样的错误,如果他们可以通过重新排序替代预测摘要减少。这导致了蕴涵模型的一个有趣的下游应用。在我们的实验中,我们发现在NLI数据集上训练的开箱即用的蕴涵模型还没有为下游任务提供所需的性能,因此我们发布了我们的注释作为额外的测试数据,用于未来NLI的外部评估。
While recent progress on abstractive summarization has led to remarkably fluent summaries, factual errors in generated summaries still severely limit their use in practice. In this paper, we evaluate summaries produced by state-of-the-art models via crowdsourcing and show that such errors occur frequently, in particular with more abstractive models. We study whether textual entailment predictions can be used to detect such errors and if they can be reduced by reranking alternative predicted summaries. That leads to an interesting downstream application for entailment models. In our experiments, we find that out-of-the-box entailment models trained on NLI datasets do not yet offer the desired performance for the downstream task and we therefore release our annotations as additional test data for future extrinsic evaluations of NLI.