Realistic Evaluation Principles for Cross-document Coreference Resolution

Realistic Evaluation Principles for Cross-document Coreference Resolution
复制标题

跨文档共指消解的现实评估原则

DOI:
10.18653/v1/2021.starsem-1.13
复制
发表时间:
2021
期刊:
ArXiv
影响因子:
--
通讯作者:
Ido Dagan
Ido Dagan
中科院分区:
--
文献类型:
--
作者:
Arie Cattan;Alon Eirew;Gabriel Stanovsky;Mandar Joshi;Ido Dagan

文献摘要

被引文献

相似文献

我们指出,跨文档共指消解的常见评估实践在其假设的设置中是不切实际的,产生了夸大的结果。我们建议通过两个评价方法原则来解决这个问题。首先,与其他任务一样,模型应该根据预测提及而不是黄金提及进行评估。这样做提出了一个微妙的问题,关于单例共指集群,我们通过解耦的评价提到检测共指链接。其次,我们认为模型不应该利用标准ECB+数据集的合成主题结构,迫使模型面对词汇歧义的挑战,正如数据集创建者所期望的那样。我们实证证明了我们更现实的评价原则对竞争模型的巨大影响,得到的分数比以前宽松的做法评估低33 F1。
We point out that common evaluation practices for cross-document coreference resolution have been unrealistically permissive in their assumed settings, yielding inflated results. We propose addressing this issue via two evaluation methodology principles. First, as in other tasks, models should be evaluated on predicted mentions rather than on gold mentions. Doing this raises a subtle issue regarding singleton coreference clusters, which we address by decoupling the evaluation of mention detection from that of coreference linking. Second, we argue that models should not exploit the synthetic topic structure of the standard ECB+ dataset, forcing models to confront the lexical ambiguity challenge, as intended by the dataset creators. We demonstrate empirically the drastic impact of our more realistic evaluation principles on a competitive model, yielding a score which is 33 F1 lower compared to evaluating by prior lenient practices.