A critique and improvement of an evaluation metric for text segmentation

A critique and improvement of an evaluation metric for text segmentation
复制标题

DOI:
10.1162/089120102317341756
复制
发表时间:
2002-03-01
影响因子:
9.3
通讯作者:
Hearst, MA
Hearst, MA
中科院分区:
计算机科学3区
文献类型:
--
作者:
Pevzner, L;Hearst, MA

文献摘要

被引文献

相似文献

最早由Beeferman,Berger和Lafferty(1997)提出的P-k评价度量正在成为评估文本分割算法的标准度量。然而,对该指标的理论分析发现了几个问题:该指标对假阴性的惩罚比对假阳性的惩罚更严重,对近距离未命中的惩罚过度,并且受到段大小分布变化的影响。我们对P-k度量提出了一个简单的修改来解决这些问题。这种名为WindowDiff的新度量在文本中移动固定大小的窗口,并在窗口内的边界数量与该文本窗口的真实边界数量不匹配时惩罚该算法。
The P-k evaluation metric, initially proposed by Beeferman, Berger, and Lafferty (1997), is becoming the standard measure for assessing text segmentation algorithms. However, a theoretical analysis of the metric finds several problems: the metric penalizes false negatives more heavily than false positives, overpenalizes near misses, and is affected by variation in segment size distribution. We propose a simple modification to the P-k metric that remedies these problems. This new metric-called WindowDiff-moves a fixed-sized window across the text and penalizes the algorithm whenever the number of boundaries within the window does not match the true number of boundaries for that window of text.