Keep Meeting Summaries on Topic: Abstractive Multi-Modal Meeting Summarization

Keep Meeting Summaries on Topic: Abstractive Multi-Modal Meeting Summarization
复制标题

DOI:
10.18653/v1/p19-1210
复制
发表时间:
2019-07
期刊:
--
影响因子:
--
通讯作者:
Manling Li;Lingyu Zhang;Heng Ji;R. Radke
Manling Li;Lingyu Zhang;Heng Ji;R. Radke
中科院分区:
其他
文献类型:
--
作者:
Manling Li;Lingyu Zhang;Heng Ji;R. Radke

文献摘要

被引文献

相似文献

自然的多人会议的文字记录与新闻文章等文档有很大的不同,这可能会使用于生成摘要的自然语言生成模型失去重点。我们开发了一个抽象的会议摘要从视频和音频的会议录音。具体来说,我们提出了一个多通道的分层注意三个层次:段,话语和单词。为了将焦点缩小到主题相关的部分,我们联合建模主题分割和摘要。除了传统的文本功能,我们引入了新的多模态功能来自视觉焦点的注意,基于这样的假设,即话语是更重要的,如果扬声器得到更多的关注。实验表明,我们的模型显着优于国家的最先进的BLEU和ROUGE措施。
Transcripts of natural, multi-person meetings differ significantly from documents like news articles, which can make Natural Language Generation models for generating summaries unfocused. We develop an abstractive meeting summarizer from both videos and audios of meeting recordings. Specifically, we propose a multi-modal hierarchical attention across three levels: segment, utterance and word. To narrow down the focus into topically-relevant segments, we jointly model topic segmentation and summarization. In addition to traditional text features, we introduce new multi-modal features derived from visual focus of attention, based on the assumption that the utterance is more important if the speaker receives more attention. Experiments show that our model significantly outperforms the state-of-the-art with both BLEU and ROUGE measures.