Evaluating rhyme annotations for large corpora Metrics and data

Evaluating rhyme annotations for large corpora Metrics and data
复制标题

评估大型语料库的韵律注释 指标和数据

DOI:
10.1163/19606028-bja10032
复制
发表时间:
2023
影响因子:
--
通讯作者:
BALEY J
BALEY J
中科院分区:
--
文献类型:
--
作者:
BALEY J

文献摘要

相似文献

针对大型押韵语料库的自动押韵标注,近年来提出了一些新的方法。这些方法,如Baley (2022b)大大降低了注释押韵材料的成本,使历史语言学家能够专注于押韵模式的分析。然而,这些注释质量的证据是轶事,由少数个别诗歌案例研究组成。本文提出了解决这个问题的方法:首先,我们讨论了之前提出的评估注释者输出质量的标准(List, Hill, and Foster; 2019),并提出了一个更适合该任务的替代指标。然后,从Baley发表的带注释的语料库中采样并手工重新注释,我们使用样本来演示原始方法中的漏洞并展示如何修复它们。最后,手工注释的示例和源代码作为附加数据发布,以便其他研究人员可以比较他们自己的注释器的性能。
Recent methods have been proposed to produce automatic rhyme annotators for large rhymed corpora. These methods, such as Baley (2022b) greatly reduce the cost of annotating rhymed material, allowing historical linguists to focus on the analysis of the rhyme patterns. However, evidence for the quality of those annotations has been anecdotal, consisting of a handful of individual poem case studies. This paper proposes to address the issue: first, we discuss previously proposed metrics that evaluate the quality of an annotator’s output against a ground-truth annotation (List, Hill, and Foster; 2019) and we propose an alternative metric that is better suited to the task. Then, sampling from Baley’s published annotated corpus and re-annotating it by hand, we use the sample to demonstrate the lacunae in the original approach and show how to fix them. Finally, the hand-annotated sample and source code are published as additional data, so that other researchers can compare the performance of their own annotators.