CATS: Customizable Abstractive Topic-based Summarization

CATS: Customizable Abstractive Topic-based Summarization
复制标题

DOI:
10.1145/3464299
复制
发表时间:
2021-10
期刊:
ACM Transactions on Information Systems (TOIS)
影响因子:
--
通讯作者:
Seyed Ali Bahrainian;George Zerveas;F. Crestani;Carsten Eickhoff
Seyed Ali Bahrainian;George Zerveas;F. Crestani;Carsten Eickhoff
中科院分区:
其他
文献类型:
--
作者:
Seyed Ali Bahrainian;George Zerveas;F. Crestani;Carsten Eickhoff

文献摘要

相似文献

神经序列到序列模型是在文本文档的抽象摘要中使用的最先进的方法,用于产生源文本叙述的精简版本,而不限于仅使用原始文本中的单词。尽管在抽象摘要方面取得了进步,但是摘要的定制生成(例如,用户的偏好)仍然未被探索。在这篇文章中,我们提出了CATS,一个抽象的神经摘要模型,以序列到序列的方式总结内容,同时还引入了一种新的机制来控制所产生的摘要的潜在主题分布。我们经验性地说明了我们的模型在生产定制的摘要和目前的研究结果,有利于设计这样的系统的有效性。我们使用著名的CNN/DailyMail数据集来评估我们的模型。此外,我们提出了一种迁移学习方法,并证明了我们的方法在低资源环境中的有效性,即,会议纪要的抽象摘要,其中结合了主要可用的会议记录数据集、AMI和国际计算机科学研究所(ICSI),仅产生几百个培训文档。
Neural sequence-to-sequence models are the state-of-the-art approach used in abstractive summarization of textual documents, useful for producing condensed versions of source text narratives without being restricted to using only words from the original text. Despite the advances in abstractive summarization, custom generation of summaries (e.g., towards a user’s preference) remains unexplored. In this article, we present CATS, an abstractive neural summarization model that summarizes content in a sequence-to-sequence fashion while also introducing a new mechanism to control the underlying latent topic distribution of the produced summaries. We empirically illustrate the efficacy of our model in producing customized summaries and present findings that facilitate the design of such systems. We use the well-known CNN/DailyMail dataset to evaluate our model. Furthermore, we present a transfer-learning method and demonstrate the effectiveness of our approach in a low resource setting, i.e., abstractive summarization of meetings minutes, where combining the main available meetings’ transcripts datasets, AMI and International Computer Science Institute(ICSI), results in merely a few hundred training documents.