A multi-level methodology for the automated translation of a coreference resolution dataset: an application to the Italian language

A multi-level methodology for the automated translation of a coreference resolution dataset: an application to the Italian language
复制标题

DOI:
10.1007/s00521-022-07641-3
复制
发表时间:
2022-09
影响因子:
6
通讯作者:
Aniello Minutolo;R. Guarasci;Emanuele Damiano;G. De Pietro;Hamido Fujita;M. Esposito
Aniello Minutolo;R. Guarasci;Emanuele Damiano;G. De Pietro;Hamido Fujita;M. Esposito
中科院分区:
计算机科学3区
文献类型:
--
作者:
Aniello Minutolo;R. Guarasci;Emanuele Damiano;G. De Pietro;Hamido Fujita;M. Esposito

文献摘要

相似文献

在过去的十年中,对易于访问的语料库的需求已经触及自然语言处理的所有领域,包括共指消解。然而,它是最近发展中考虑最少的子领域之一。此外,几乎所有现有资源都只提供英文。为了克服这一不足,这项工作提出了一种方法来创建一个语料库共指消解意大利语利用知识的注释资源在其他语言。从OntonNotes开始,该方法翻译和精炼英语话语,以获得尊重意大利语语法,处理语言特有的现象和保留共指和提及的话语。通过定量和定性的评价,综合考虑可读性、语法性和可接受性等指标,对生成的话语进行形式良好性的评估。结果证实了该方法在从现有数据集开始生成用于共指消解的良好数据集方面的有效性。通过训练一个基于BERT语言模型的共指消解模型来评估数据集的优度,取得了令人满意的结果。即使该方法是为英语和意大利语量身定制的,它也有一个通用的基础,可以很容易地扩展到其他语言,采用少量的语言相关规则来概括所研究语言的大多数语言现象。
In the last decade, the demand for readily accessible corpora has touched all areas of natural language processing, including coreference resolution. However, it is one of the least considered sub-fields in recent developments. Moreover, almost all existing resources are only available for the English language. To overcome this lack, this work proposes a methodology to create a corpus for coreference resolution in Italian exploiting knowledge of annotated resources in other languages. Starting from OntonNotes, the methodology translates and refines English utterances to obtain utterances respecting Italian grammar, dealing with language-specific phenomena and preserving coreference and mentions. A quantitative and qualitative evaluation is performed to assess the well-formedness of generated utterances, considering readability, grammaticality, and acceptability indexes. The results have confirmed the effectiveness of the methodology in generating a good dataset for coreference resolution starting from an existing one. The goodness of the dataset is also assessed by training a coreference resolution model based on BERT language model, achieving the promising results. Even if the methodology has been tailored for English and Italian languages, it has a general basis easily extendable to other languages, adapting a small number of language-dependent rules to generalize most of the linguistic phenomena of the language under examination.