MediaSum: A Large-scale Media Interview Dataset for Dialogue Summarization

MediaSum: A Large-scale Media Interview Dataset for Dialogue Summarization
复制标题

DOI:
10.18653/v1/2021.naacl-main.474
复制
发表时间:
2021-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Chenguang Zhu;Yang Liu;Jie Mei;Michael Zeng
Chenguang Zhu;Yang Liu;Jie Mei;Michael Zeng
中科院分区:
其他
文献类型:
--
作者:
Chenguang Zhu;Yang Liu;Jie Mei;Michael Zeng

文献摘要

被引文献

相似文献

MediaSum是一个大型媒体采访数据集,包含463.6K文本和抽象摘要。为了创建这个数据集,我们收集了来自NPR和CNN的采访记录,并使用概述和主题描述作为摘要。与现有的用于对话摘要的公共语料库相比,我们的数据集大了一个数量级,并且包含来自多个领域的复杂多方对话。我们进行统计分析,以证明独特的位置偏见表现在电视和广播采访的文字记录。我们还表明,MediaSum可以用于迁移学习,以提高模型在其他对话摘要任务上的表现。
This paper introduces MediaSum, a large-scale media interview dataset consisting of 463.6K transcripts with abstractive summaries. To create this dataset, we collect interview transcripts from NPR and CNN and employ the overview and topic descriptions as summaries. Compared with existing public corpora for dialogue summarization, our dataset is an order of magnitude larger and contains complex multi-party conversations from multiple domains. We conduct statistical analysis to demonstrate the unique positional bias exhibited in the transcripts of televised and radioed interviews. We also show that MediaSum can be used in transfer learning to improve a model’s performance on other dialogue summarization tasks.