Diachronic proximity vs. data sparsity in cross-lingual parser projection. A case study on Germanic

Diachronic proximity vs. data sparsity in cross-lingual parser projection. A case study on Germanic
复制标题

跨语言解析器投影中的历时邻近性与数据稀疏性。

DOI:
10.3115/v1/w14-5302
复制
发表时间:
2014
影响因子:
3
通讯作者:
C. Chiarcos
C. Chiarcos
中科院分区:
工程技术3区
文献类型:
--
作者:
Maria Sukhareva;C. Chiarcos

文献摘要

被引文献

相似文献

对于历史语言变体的研究,训练数据的稀疏性给句法注释和自动化过程的NLP工具的开发带来了巨大的问题。在本文中,我们探索的策略,以弥补缺乏训练数据,包括相关品种的数据,在一系列的注释投影实验从英语到四个古老的日耳曼语言:在依赖句法投影从英语到一个或多个语言(S),我们训练片段感知的解析器训练,并将其应用到目标语言。对于解析器训练,我们将目标语言的小数据集作为基线,并将其与在来自不同相关程度的多个品种的较大数据集上训练的模型进行比较,从而平衡稀疏性和历时接近性。我们的实验表明:(a)在目标语言的训练数据中包含相关语言数据可以提高句法分析性能,
For the study of historical language varieties, the sparsity of training data imposes immense problems on syntactic annotation and the development of NLP tools that automatize the process. In this paper, we explore strategies to compensate the lack of training data by including data from related varieties in a series of annotation projection experiments from English to four old Germanic languages: On dependency syntax projected from English to one or multiple language(s), we train a fragment-aware parser trained and apply it to the target language. For parser training, we consider small datasets from the target language as a baseline, and compare it with models trained on larger datasets from multiple varieties with different degrees of relatedness, thereby balancing sparsity and diachronic proximity. Our experiments show (a) that including related language data to training data in the target language can improve parsing performance,