Abstractive Document Summarization with Word Embedding Reconstruction
Abstractive Document Summarization with Word Embedding Reconstruction
复制标题
DOI:
10.26615/978-954-452-072-4_178
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Jingyi You;Chenlong Hu;Hidetaka Kamigaito;Hiroya Takamura;M. Okumura
中科院分区:
文献类型:
--
作者:
Jingyi You;Chenlong Hu;Hidetaka Kamigaito;Hiroya Takamura;M. Okumura
Neural sequence-to-sequence (Seq2Seq) models and BERT have achieved substantial improvements in abstractive document summarization (ADS) without and with pre-training, respectively. However, they sometimes repeatedly attend to unimportant source phrases while mistakenly ignore important ones. We present reconstruction mechanisms on two levels to alleviate this issue. The sequence-level reconstructor reconstructs the whole document from the hidden layer of the target summary, while the word embedding-level one rebuilds the average of word embeddings of the source at the target side to guarantee that as much critical information is included in the summary as possible. Based on the assumption that inverse document frequency (IDF) measures how important a word is, we further leverage the IDF weights in our embedding-level reconstructor. The proposed frameworks lead to promising improvements for ROUGE metrics and human rating on both the CNN/Daily Mail and Newsroom summarization datasets.