Keep Meeting Summaries on Topic: Abstractive Multi-Modal Meeting Summarization
Keep Meeting Summaries on Topic: Abstractive Multi-Modal Meeting Summarization
复制标题
DOI:
10.18653/v1/p19-1210
复制
发表时间:
2019-07
期刊:
影响因子:
--
通讯作者:
Manling Li;Lingyu Zhang;Heng Ji;R. Radke
中科院分区:
文献类型:
--
作者:
Manling Li;Lingyu Zhang;Heng Ji;R. Radke
Transcripts of natural, multi-person meetings differ significantly from documents like news articles, which can make Natural Language Generation models for generating summaries unfocused. We develop an abstractive meeting summarizer from both videos and audios of meeting recordings. Specifically, we propose a multi-modal hierarchical attention across three levels: segment, utterance and word. To narrow down the focus into topically-relevant segments, we jointly model topic segmentation and summarization. In addition to traditional text features, we introduce new multi-modal features derived from visual focus of attention, based on the assumption that the utterance is more important if the speaker receives more attention. Experiments show that our model significantly outperforms the state-of-the-art with both BLEU and ROUGE measures.