Formal and functional assessment of the pyramid method for summary content evaluation*

Formal and functional assessment of the pyramid method for summary content evaluation*
复制标题

DOI:
10.1017/s1351324909005051
复制
发表时间:
2009-04
影响因子:
2.5
通讯作者:
R. Passonneau
R. Passonneau
中科院分区:
计算机科学3区
文献类型:
--
作者:
R. Passonneau

文献摘要

被引文献

相似文献

摘要金字塔注释使得定量和定性地评估机器生成(或人类)摘要的内容成为可能。评价方法必须与其他研究方法一样,以同样的衡量标准-评价-来证明自己。首先,从2003年的文件理解会议(DUC)的金字塔数据的正式评估,这解决了注释的形式是否可靠,以及评分结果是否是一致的注释。两个手动注释阶段(金字塔创建和对金字塔模型的系统同行摘要的注释)的interannotator可靠性测量的组合,以及来自不同注释的系统分数的相似性的显著性测试,产生高度可靠的结果。最严格的测试包括比较两组独立的金字塔和同行注释产生的同行系统排名,它们产生基本相同的排名。三年的DUC数据(2003年,2005年,2006年)用于评估不同评价环境中方法的可靠性:不同系统,文件集,总结长度和模型总结数量。该功能评估解决了该方法跨年度区分系统的能力。结果表明,该方法的统计功效足以识别系统之间的统计显著差异,并且统计功效在3年内变化不大。
Abstract Pyramid annotation makes it possible to evaluate quantitatively and qualitatively the content of machine-generated (or human) summaries. Evaluation methods must prove themselves against the same measuring stick – evaluation – as other research methods. First, a formal assessment of pyramid data from the 2003 Document Understanding Conference (DUC) is presented; this addresses whether the form of annotation is reliable and whether score results are consistent across annotators. A combination of interannotator reliability measures of the two manual annotation phases (pyramid creation and annotation of system peer summaries against pyramid models), and significance tests of the similarity of system scores from distinct annotations, produces highly reliable results. The most rigorous test consists of a comparison of peer system rankings produced from two independent sets of pyramid and peer annotations, which produce essentially the same rankings. Three years of DUC data (2003, 2005, 2006) are used to assess the reliability of the method across distinct evaluation settings: distinct systems, document sets, summary lengths, and numbers of model summaries. This functional assessment addresses the method's ability to discriminate systems across years. Results indicate that the statistical power of the method is more than sufficient to identify statistically significant differences among systems, and that the statistical power varies little across the 3 years.