Automatically Assessing Machine Summary Content Without a Gold Standard

Automatically Assessing Machine Summary Content Without a Gold Standard
复制标题

DOI:
10.1162/coli_a_00123
复制
发表时间:
2013-06
影响因子:
9.3
通讯作者:
Annie Louis;A. Nenkova
Annie Louis;A. Nenkova
中科院分区:
计算机科学3区
文献类型:
--
作者:
Annie Louis;A. Nenkova

文献摘要

被引文献

相似文献

用于评估摘要内容的最广泛采用的方法遵循一些协议,用于将摘要与黄金标准的人类摘要(传统上称为模型摘要)进行比较。当无法获得人工总结时,这种评估范式就会不足,当只有一个模型可用时,这种评估范式就会变得不那么准确。我们提出了三种新的评估技术。其中两个是无模型的,不依赖于评估的黄金标准。第三种技术通过使用选择的系统摘要扩展可用的模型摘要集来改进标准的自动评估。我们表明,量化源文本与其摘要之间的相似性与适当选择的措施产生总结分数,准确地复制人类的评估。我们还探索了在只有一个人类模型摘要作为金标准可用时提高评估质量的方法。我们引入了伪模型,这些伪模型是根据自动评估认为包含良好内容的系统摘要。与只使用一个可用模型相比,将伪模型与单个人类模型结合起来形成黄金标准,与人类判断的相关性更高。最后,我们探讨了另一种度量的可行性——相同输入的系统摘要和所有其他系统摘要池之间的相似性。这种与系统的共识进行比较的方法产生了令人印象深刻的系统摘要的精确排名,实现了与人类排名在0.9以上的相关性。
The most widely adopted approaches for evaluation of summary content follow some protocol for comparing a summary with gold-standard human summaries, which are traditionally called model summaries. This evaluation paradigm falls short when human summaries are not available and becomes less accurate when only a single model is available. We propose three novel evaluation techniques. Two of them are model-free and do not rely on a gold standard for the assessment. The third technique improves standard automatic evaluations by expanding the set of available model summaries with chosen system summaries.We show that quantifying the similarity between the source text and its summary with appropriately chosen measures produces summary scores which replicate human assessments accurately. We also explore ways of increasing evaluation quality when only one human model summary is available as a gold standard. We introduce pseudomodels, which are system summaries deemed to contain good content according to automatic evaluation. Combining the pseudomodels with the single human model to form the gold-standard leads to higher correlations with human judgments compared to using only the one available model. Finally, we explore the feasibility of another measure—similarity between a system summary and the pool of all other system summaries for the same input. This method of comparison with the consensus of systems produces impressively accurate rankings of system summaries, achieving correlation with human rankings above 0.9.