NEWTS: A Corpus for News Topic-Focused Summarization

NEWTS: A Corpus for News Topic-Focused Summarization
复制标题

DOI:
10.48550/arxiv.2205.15661
复制
发表时间:
2022-05
期刊:
--
影响因子:
--
通讯作者:
Seyed Ali Bahrainian;Sheridan Feucht;Carsten Eickhoff
Seyed Ali Bahrainian;Sheridan Feucht;Carsten Eickhoff
中科院分区:
其他
文献类型:
--
作者:
Seyed Ali Bahrainian;Sheridan Feucht;Carsten Eickhoff

文献摘要

被引文献

相似文献

文本摘要模型正在接近人类的保真度水平。现有的基准语料库提供了完整和删节版本的网络,新闻或专业内容的一致对。到目前为止,所有的摘要数据集都是在一种通用的模式下运行的,这种模式可能无法反映有机摘要的全部需求。最近提出的几种模型(例如,即插即用语言模型)具有根据期望的主题范围来调节所生成的概要的能力。这些能力仍然在很大程度上未被使用和未评估,因为没有专门的数据集,将支持的主题为重点的summarization.This任务本文介绍了第一个主题摘要语料库NEWTS,基于著名的CNN/Dailymail数据集,并通过在线众包注释。每一篇源文章都有两个参考摘要,每个参考摘要都关注源文档的不同主题。我们评估了一系列有代表性的现有技术,并分析了不同的激励方法的有效性。
Text summarization models are approaching human levels of fidelity. Existing benchmarking corpora provide concordant pairs of full and abridged versions of Web, news or professional content. To date, all summarization datasets operate under a one-size-fits-all paradigm that may not reflect the full range of organic summarization needs. Several recently proposed models (e.g., plug and play language models) have the capacity to condition the generated summaries on a desired range of themes. These capacities remain largely unused and unevaluated as there is no dedicated dataset that would support the task of topic-focused summarization.This paper introduces the first topical summarization corpus NEWTS, based on the well-known CNN/Dailymail dataset, and annotated via online crowd-sourcing. Each source article is paired with two reference summaries, each focusing on a different theme of the source document. We evaluate a representative range of existing techniques and analyze the effectiveness of different prompting methods.