Exploring Correlation Between ROUGE and Human Evaluation on Meeting Summaries

Exploring Correlation Between ROUGE and Human Evaluation on Meeting Summaries
复制标题

探索 ROUGE 与人类对会议摘要评估之间的相关性

DOI:
10.1109/tasl.2009.2025096
复制
发表时间:
2010
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Yang Liu
Yang Liu
中科院分区:
--
文献类型:
--
作者:
F. Liu;Yang Liu

文献摘要

被引文献

相似文献

摘要的自动评测是文摘系统开发的一个重要环节。在文本摘要中,ROUGE已经被证明在测量内容单元的匹配时与人类评价很好地相关。然而,多方会议域有许多特点,这可能会给ROUGE带来潜在的问题。本文的目标是研究如何以及ROUGE分数与人类评价提取会议摘要,并探讨不同的会议领域的具体因素,有影响的相关性。在本研究中,我们进行了更多的分析,比我们以前的工作。我们的实验表明,一般ROUGE和人的评价之间的相关性不是很大,但是,当占几个独特的会议特征,如不流利,扬声器信息,和停用词在ROUGE设置,可以实现更好的相关性,特别是在系统摘要。我们还发现,这些因素对人类与系统摘要有不同的影响。此外,我们对比使用ROUGE与其他自动摘要评价指标,如Kappa和金字塔的结果,并显示使用ROUGE的适当性。
Automatic summarization evaluation is very important to the development of summarization systems. In text summarization, ROUGE has been shown to correlate well with human evaluation when measuring match of content units. However, there are many characteristics of the multiparty meeting domain, which may pose potential problems to ROUGE. The goal of this paper is to examine how well the ROUGE scores correlate with human evaluation for extractive meeting summarization, and explore different meeting domain specific factors that have an impact on the correlation. More analysis than those in our previous work has been conducted in this study. Our experiments show that generally the correlation between ROUGE and human evaluation is not great; however, when accounting for several unique meeting characteristics, such as disfluencies, speaker information, and stopwords in the ROUGE setting, better correlation can be achieved, especially on the system summaries. We also found that these factors have a different impact on human versus system summaries. In addition, we contrast the results using ROUGE with other automatic summarization evaluation metrics, such as Kappa and Pyramid, and show the appropriateness of using ROUGE for this study.