Cross-Lingual Morphological Tagging for Low-Resource Languages

Cross-Lingual Morphological Tagging for Low-Resource Languages
复制标题

低资源语言的跨语言形态标记

DOI:
10.18653/v1/p16-1184
复制
发表时间:
2016
期刊:
ArXiv
影响因子:
--
通讯作者:
Jan A. Botha
Jan A. Botha
中科院分区:
--
文献类型:
--
作者:
Jan Buys;Jan A. Botha

文献摘要

被引文献

相似文献

形态丰富的语言通常缺乏开发准确的自然语言处理工具所需的注释语言资源。我们提出了适合在不使用直接监督的情况下为低资源语言训练具有丰富标记集的形态标记器的模型。我们的方法扩展了跨语言投影词性标签的现有方法,使用双文本来推断对给定单词类型或标记的可能标签的约束。我们提出了一种使用 Wsabie 的标记模型,这是一种具有基于排名学习的判别性嵌入模型。在我们对 11 种语言的评估中,该模型的平均性能与基线弱监督 HMM 相当,同时更具可扩展性。多语言实验表明,该方法在相关语言对之间进行投影时表现最佳。尽管存在固有的有损投影,但我们表明,我们的模型预测的形态标签将解析器的下游性能平均提高了 +0.6 LAS。
Morphologically rich languages often lack the annotated linguistic resources required to develop accurate natural language processing tools. We propose models suitable for training morphological taggers with rich tagsets for low-resource languages without using direct supervision. Our approach extends existing approaches of projecting part-of-speech tags across languages, using bitext to infer constraints on the possible tags for a given word type or token. We propose a tagging model using Wsabie, a discriminative embeddingbased model with rank-based learning. In our evaluation on 11 languages, on average this model performs on par with a baseline weakly-supervised HMM, while being more scalable. Multilingual experiments show that the method performs best when projecting between related language pairs. Despite the inherently lossy projection, we show that the morphological tags predicted by our models improve the downstream performance of a parser by +0.6 LAS on average.