Optimization of Reference-less Evaluation Metric of Grammatical Error Correction for Manual Evaluations

Optimization of Reference-less Evaluation Metric of Grammatical Error Correction for Manual Evaluations
复制标题

DOI:
10.5715/jnlp.28.404
复制
发表时间:
2021
期刊:
Journal of Natural Language Processing
影响因子:
--
通讯作者:
Ryoma Yoshimura;Masahiro Kaneko;Tomoyuki Kajiwara;Mamoru Komachi
Ryoma Yoshimura;Masahiro Kaneko;Tomoyuki Kajiwara;Mamoru Komachi
中科院分区:
其他
文献类型:
--
作者:
Ryoma Yoshimura;Masahiro Kaneko;Tomoyuki Kajiwara;Mamoru Komachi

文献摘要

相似文献

建立一个可靠的语法纠错自动评价指标,对语法纠错的研究和发展具有重要意义。由于覆盖所有可能的参考句是困难的,以前的研究提出了无参考度量。其中之一实现了更高的相关性,人工评估比基于参考的指标,从语法,简洁性和意义保存的三个角度整合指标。然而,可以进一步改进与手动评估的相关性,因为它们不被考虑用于优化每个手动评估的每个度量。因此,在这项研究中,我们提出了一种优化每个指标的方法。此外,我们创建了一个数据集,对系统输出进行手动评估,这是优化的理想选择。实验结果表明,该方法在各个视角的度量和度量的组合上都提高了与人工评价的相关性。我们还证明了使用预训练的语言模型进行优化和优化GEC系统输出的手动评估都有助于改进。作为分析的结果,我们发现,我们提出的指标适当奖励更多的错误类型的编辑比传统的方法。
The development of a reliable automatic evaluation metric of grammatical error correction (GEC) is useful for the research and development of GEC. Since it is difficult to cover all possible reference sentences, previous studies have proposed reference-less metrics. One of them achieved a higher correlation with manual evaluation than reference-based metrics by integrating metrics from the three perspectives of grammaticality, fluency, and meaning preservation. However, the correlation with the manual evaluation can be further improved because they are not considered for optimizing each metric for each manual evaluation. Therefore, in this study, we propose a method of optimizing each metric. Furthermore, we create a dataset with manual evaluation of system output that is ideal for optimization. Experimental results show that the proposed method improves correlation with the manual evaluation in both the metric of each perspective and combining the metrics. We also demonstrate that both using pre-trained language models for optimization and optimizing to manual evaluation on system output of GEC contribute to improvement. As a result of the analysis, it was found that our proposed metric appropriately rewarded more edits of error types than the conventional methods.