Diachronic proximity vs. data sparsity in cross-lingual parser projection. A case study on Germanic
Diachronic proximity vs. data sparsity in cross-lingual parser projection. A case study on Germanic
复制标题
跨语言解析器投影中的历时邻近性与数据稀疏性。
DOI:
10.3115/v1/w14-5302
复制
发表时间:
2014
影响因子:
3
通讯作者:
C. Chiarcos
中科院分区:
文献类型:
--
作者:
Maria Sukhareva;C. Chiarcos
For the study of historical language varieties, the sparsity of training data imposes immense problems on syntactic annotation and the development of NLP tools that automatize the process. In this paper, we explore strategies to compensate the lack of training data by including data from related varieties in a series of annotation projection experiments from English to four old Germanic languages: On dependency syntax projected from English to one or multiple language(s), we train a fragment-aware parser trained and apply it to the target language. For parser training, we consider small datasets from the target language as a baseline, and compare it with models trained on larger datasets from multiple varieties with different degrees of relatedness, thereby balancing sparsity and diachronic proximity. Our experiments show (a) that including related language data to training data in the target language can improve parsing performance,