Evaluating the Efficacy of Summarization Evaluation across Languages

Evaluating the Efficacy of Summarization Evaluation across Languages
复制标题

DOI:
10.18653/v1/2021.findings-acl.71
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Fajri Koto;Jey Han Lau;Timothy Baldwin
Fajri Koto;Jey Han Lau;Timothy Baldwin
中科院分区:
其他
文献类型:
--
作者:
Fajri Koto;Jey Han Lau;Timothy Baldwin

文献摘要

被引文献

相似文献

虽然为英语开发的自动摘要评估方法通常应用于其他语言,但这是首次尝试系统地量化其泛语言功效。我们采用八种不同语言的摘要语料库,并手动注释生成的摘要,以获得焦点(精确度)和覆盖率(召回率)。在此基础上,我们评估了19个摘要评估指标,发现在BERTScore中使用多语言BERT在所有语言中表现良好,高于英语。
While automatic summarization evaluation methods developed for English are routinely applied to other languages, this is the first attempt to systematically quantify their panlinguistic efficacy. We take a summarization corpus for eight different languages, and manually annotate generated summaries for focus (precision) and coverage (recall). Based on this, we evaluate 19 summarization evaluation metrics, and find that using multilingual BERT within BERTScore performs well across all languages, at a level above that for English.