Two-Encoder Pointer-Generator Network for Summarizing Segments of Long Articles

Two-Encoder Pointer-Generator Network for Summarizing Segments of Long Articles
复制标题

DOI:
10.1007/978-3-030-26072-9_23
复制
发表时间:
2019-08
期刊:
--
影响因子:
--
通讯作者:
Junhao Li;M. Iwaihara
Junhao Li;M. Iwaihara
中科院分区:
其他
文献类型:
--
作者:
Junhao Li;M. Iwaihara

文献摘要

相似文献

通常,长文档包含许多节和段。在维基百科中,一个条目通常可以分为几个部分,一个部分可以分为几个部分。但是,尽管一篇文章已经被分成了更小的部分,一个部分仍然可能太长而无法阅读。因此,我们认为片段应该有一个简短的摘要,以便读者快速了解片段。本文讨论了将Seq2Seq模型和指针生成器网络模型等神经摘要模型应用于分段摘要。这些用于摘要的模型可以将目标段作为模型的唯一输入。然而,在我们的例子中,同一篇文章中的其余部分很可能包含与目标部分相关的描述。因此,我们提出了几种方法来提取一个额外的序列从整个文章,然后联合收割机与目标段,提供作为输入的摘要。我们将结果与没有额外序列的原始模型进行比较。此外,我们提出了一个新的模型,使用两个编码器分别处理目标段和附加序列。我们的结果表明,我们的两个编码器模型优于原始模型的ROGUE和METEOR分数。
Usually long documents contain many sections and segments. In Wikipedia, one article can usually be divided into sections and one section can be divided into segments. But although one article is already divided into smaller segments, one segment can still be too long to read. So, we consider that segments should have a short summary for readers to grasp a quick view of the segment. This paper discusses applying neural summarization models including Seq2Seq model and pointer generator network model to segment summarization. These models for summarization can take target segments as the only input to the model. However, in our case, it is very likely that the remaining segments in the same article contain descriptions related to the target segment. Therefore, we propose several ways to extract an additional sequence from the whole article and then combine with the target segment, to be supplied as the input for summarization. We compare the results against the original models without additional sequences. Furthermore, we propose a new model that uses two encoders to process the target segment and additional sequence separately. Our results show our two-encoder model outperforms the original models in terms of ROGUE and METEOR scores.