URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors
URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors
复制标题
URIEL 和 lang2vec:将语言表示为类型学、地理和系统发育向量
DOI:
10.18653/v1/e17-2002
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
Carlisle Turner
中科院分区:
文献类型:
--
作者:
Lori S. Levin;Patrick Littell;David R. Mortensen;Ke Lin;Katherine Kairis;Carlisle Turner
We introduce the URIEL knowledge base for massively multilingual NLP and the lang2vec utility, which provides information-rich vector identifications of languages drawn from typological, geographical, and phylogenetic databases and normalized to have straightforward and consistent formats, naming, and semantics. The goal of URIEL and lang2vec is to enable multilingual NLP, especially on less-resourced languages and make possible types of experiments (especially but not exclusively related to NLP tasks) that are otherwise difficult or impossible due to the sparsity and incommensurability of the data sources. lang2vec vectors have been shown to reduce perplexity in multilingual language modeling, when compared to one-hot language identification vectors.