Evaluating Information Content by Factoid Analysis: Human annotation and stability

Evaluating Information Content by Factoid Analysis: Human annotation and stability
复制标题

通过事实分析评估信息内容:人工注释和稳定性

DOI:
--
复制
发表时间:
2004
期刊:
Conference on Empirical Methods in Natural Language Processing
影响因子:
--
通讯作者:
H. Halteren
H. Halteren
中科院分区:
--
文献类型:
--
作者:
Simone Teufel;H. Halteren

文献摘要

被引文献

相似文献

基于货车Halteren和Teufel(2003)的初步实验,我们提出了一种新的内在概括性评价方法,该方法结合了两个新颖的方面:信息内容的比较(而不是字符串相似性)在黄金标准和系统摘要中,以共享的原子信息单位(我们称之为factoid)测量,并与多个黄金标准摘要进行比较(在我们的数据中:分别为20和50个总结)。在本文中,我们表明,事实的注释是高度可再生的,引入加权的事实得分,估计有多少摘要需要稳定的系统排名,并表明,事实得分不能succiently近似的unigrams和DUC信息重叠措施。
We present a new approach to intrinsic sum-mary evaluation, based on initial experiments in van Halteren and Teufel (2003), which combines two novel aspects: comparison of information content (rather than string similarity) in gold standard and system summary, measured in shared atomic information units which we call factoids , and comparison to more than one gold standard summary (in our data: 20 and 50 summaries respectively). In this paper, we show that factoid annotation is highly re-producible, introduce a weighted factoid score, estimate how many summaries are required for stable system rankings, and show that the factoid scores cannot be su–ciently approximated by unigrams and the DUC information overlap measure.