Continuous multilinguality with language vectors

Continuous multilinguality with language vectors
复制标题

语言向量的连续多语言性

DOI:
--
复制
发表时间:
2016
期刊:
Conference of the European Chapter of the Association for Computational Linguistics
影响因子:
--
通讯作者:
J. Tiedemann
J. Tiedemann
中科院分区:
--
文献类型:
--
作者:
Robert Östling;J. Tiedemann

文献摘要

被引文献

相似文献

大多数现有的多语言自然语言处理(NLP)模型将语言视为一个离散的类别,并对一种语言或另一种语言进行预测。相反,我们建议使用语言的连续向量表示。我们表明,这些可以通过基于字符的神经语言模型有效地学习,并用于改进对训练中未见的语言变体的推断。在1303个圣经翻译成990种不同语言的实验中,我们实证地探索了多语言语言模型的能力,并表明语言向量捕获了语言之间的遗传关系。
Most existing models for multilingual natural language processing (NLP) treat language as a discrete category, and make predictions for either one language or the other. In contrast, we propose using continuous vector representations of language. We show that these can be learned efficiently with a character-based neural language model, and used to improve inference about language varieties not seen during training. In experiments with 1303 Bible translations into 990 different languages, we empirically explore the capacity of multilingual language models, and also show that the language vectors capture genetic relationships between languages.