Transforming Wikipedia Into Augmented Data for Query-Focused Summarization

Transforming Wikipedia Into Augmented Data for Query-Focused Summarization
复制标题

DOI:
10.1109/taslp.2022.3171963
复制
发表时间:
2019-11
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Haichao Zhu;Li Dong;Furu Wei;Bing Qin;Ting Liu
Haichao Zhu;Li Dong;Furu Wei;Bing Qin;Ting Liu
中科院分区:
其他
文献类型:
--
作者:
Haichao Zhu;Li Dong;Furu Wei;Bing Qin;Ting Liu

文献摘要

相似文献

现有以查询为中心的摘要数据集的规模有限,使得训练数据驱动的摘要模型具有挑战性。同时,手工构建以查询为中心的摘要语料库成本高,耗时长。在本文中,我们使用Wikipedia自动收集了一个超过280,000个示例的以查询为中心的大型摘要数据集(名为WikiRef),这可以作为数据增强的一种手段。我们还开发了一个基于bert的以查询为中心的摘要模型(Q-BERT),用于从文档中提取句子作为摘要。为了更好地使包含数百万个参数的巨大模型适应微小的基准测试,我们只识别和微调一个稀疏的子网络,它对应于整个模型参数的一小部分。在三个DUC基准上的实验结果表明,在WikiRef上预训练的模型已经达到了合理的性能。在对特定基准数据集进行微调后,具有数据增强的模型优于强比较系统。此外,我们提出的Q-BERT模型和子网微调都进一步提高了模型的性能。
The limited size of existing query-focused summarization datasets renders training data-driven summarization models challenging. Meanwhile, the manual construction of a query-focused summarization corpus is costly and time-consuming. In this paper, we use Wikipedia to automatically collect a large query-focused summarization dataset (named WikiRef) of more than 280,000 examples, which can serve as a means of data augmentation. We also develop a BERT-based query-focused summarization model (Q-BERT) to extract sentences from the documents as summaries. To better adapt a huge model containing millions of parameters to tiny benchmarks, we identify and fine-tune only a sparse subnetwork, which corresponds to a small fraction of the whole model parameters. Experimental results on three DUC benchmarks show that the model pre-trained on WikiRef has already achieved reasonable performance. After fine-tuning on the specific benchmark datasets, the model with data augmentation outperforms strong comparison systems. Moreover, both our proposed Q-BERT model and subnetwork fine-tuning further improve the model performance.