A critique and improvement of an evaluation metric for text segmentation
A critique and improvement of an evaluation metric for text segmentation
复制标题
DOI:
10.1162/089120102317341756
复制
发表时间:
2002-03-01
影响因子:
9.3
通讯作者:
Hearst, MA
中科院分区:
文献类型:
--
作者:
Pevzner, L;Hearst, MA
The P-k evaluation metric, initially proposed by Beeferman, Berger, and Lafferty (1997), is becoming the standard measure for assessing text segmentation algorithms. However, a theoretical analysis of the metric finds several problems: the metric penalizes false negatives more heavily than false positives, overpenalizes near misses, and is affected by variation in segment size distribution. We propose a simple modification to the P-k metric that remedies these problems. This new metric-called WindowDiff-moves a fixed-sized window across the text and penalizes the algorithm whenever the number of boundaries within the window does not match the true number of boundaries for that window of text.