Summarization Evaluation Methods: Experiments and Analysis

Summarization Evaluation Methods: Experiments and Analysis
复制标题

总结评价方法:实验与分析

DOI:
10.7916/d8tb1g77
复制
发表时间:
1998
期刊:
影响因子:
--
通讯作者:
Michael Elhadad
Michael Elhadad
中科院分区:
--
文献类型:
--
作者:
K. McKeown;Hongyan Jing;R. Barzilay;Michael Elhadad

文献摘要

被引文献

相似文献

有两种方法用于评估摘要系统:根据“理想”摘要对生成的摘要进行评估,以及评估摘要在多大程度上帮助人们执行任务(如informa)。加强检索。我们进行了两个大型实验来研究这两种评价方法。我们的结果表明,实验的不同参数会对系统的得分产生影响。例如,总结长度被发现会影响这两种类型的评估。对于“理想”的基于摘要的评估,准确性随着摘要长度的增加而降低,而对于基于任务的评估,信息检索任务的摘要长度和准确性似乎是随机相关的。在本文中,我们展示了该参数和其他参数如何影响评估结果,并描述了如何控制参数以产生合理的评估。
Two methods are used for evaluation of summarization systems: an evaluation of generated summaries against an "ideal" summary and evaluation of how well summaries help a person perform in a task such as informa. tion retrieval. We carried out two large experiments to study the two evaluation methods. Our results show that different parameters of an experiment can (h-amatically affect how well a system scores. For example, summary length was found to affect both types of evaluations. For the "ideal" summary based evaluation, accuracy decreases as summary length increases, while for task based evaluations summary length and accuracy on an information retrieval task appear to correlate randomly. In this paper, we show how this parameter and others can affect evaluation results and describe how parameters can be controlled to produce a sound evaluation.