SQuALITY: Building a Long-Document Summarization Dataset the Hard Way

SQuALITY: Building a Long-Document Summarization Dataset the Hard Way
复制标题

DOI:
10.48550/arxiv.2205.11465
复制
发表时间:
2022-05
期刊:
--
影响因子:
--
通讯作者:
Alex Wang;Richard Yuanzhe Pang;Angelica Chen;Jason Phang;Samuel R. Bowman
Alex Wang;Richard Yuanzhe Pang;Angelica Chen;Jason Phang;Samuel R. Bowman
中科院分区:
其他
文献类型:
--
作者:
Alex Wang;Richard Yuanzhe Pang;Angelica Chen;Jason Phang;Samuel R. Bowman

文献摘要

被引文献

相似文献

摘要数据集通常是通过收集自然产生的公共领域摘要(这几乎总是在难以处理的技术领域)或使用近似启发式方法从日常文本中提取它们来组装的,这经常产生不忠实的摘要。在这项工作中,我们转向一种较慢但更直接的方法来开发摘要基准数据:我们雇用高素质的承包商从头开始阅读故事并编写原始摘要。为了分摊阅读时间,我们为每个文档收集了五个摘要,第一个概述,随后的四个解决具体问题。我们使用该协议来收集SQuALITY,这是一个以问题为中心的摘要数据集,与多项选择数据集QuALITY建立在相同的公共领域短篇故事上(Pang等人,2021)。用最先进的摘要系统进行的实验表明,我们的数据集具有挑战性,现有的自动评估指标是质量的弱指标。
Summarization datasets are often assembled either by scraping naturally occurring public-domain summaries—which are nearly always in difficult-to-work-with technical domains—or by using approximate heuristics to extract them from everyday text—which frequently yields unfaithful summaries. In this work, we turn to a slower but more straightforward approach to developing summarization benchmark data: We hire highly-qualified contractors to read stories and write original summaries from scratch. To amortize reading time, we collect five summaries per document, with the first giving an overview and the subsequent four addressing specific questions. We use this protocol to collect SQuALITY, a dataset of question-focused summaries built on the same public-domain short stories as the multiple-choice dataset QuALITY (Pang et al., 2021). Experiments with state-of-the-art summarization systems show that our dataset is challenging and that existing automatic evaluation metrics are weak indicators of quality.