URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors

URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors
复制标题

URIEL 和 lang2vec:将语言表示为类型学、地理和系统发育向量

DOI:
10.18653/v1/e17-2002
复制
发表时间:
2017
期刊:
Current Approaches to Metaphor Analysis in Discourse
影响因子:
--
通讯作者:
Carlisle Turner
Carlisle Turner
中科院分区:
--
文献类型:
--
作者:
Lori S. Levin;Patrick Littell;David R. Mortensen;Ke Lin;Katherine Kairis;Carlisle Turner

文献摘要

被引文献

相似文献

我们引入了用于大规模多语言 NLP 的 URIEL 知识库和 lang2vec 实用程序,该实用程序提供从类型学、地理和系统发育数据库中提取的语言的信息丰富的矢量识别,并标准化为具有简单且一致的格式、命名和语义。 URIEL 和 lang2vec 的目标是实现多语言 NLP,特别是在资源较少的语言上,并使实验类型成为可能(特别是但不完全与 NLP 任务相关),否则由于数据源的稀疏性和不可通约性,这些实验是困难或不可能的。与 one-hot 语言识别向量相比,lang2vec 向量已被证明可以减少多语言语言建模中的复杂性。
We introduce the URIEL knowledge base for massively multilingual NLP and the lang2vec utility, which provides information-rich vector identifications of languages drawn from typological, geographical, and phylogenetic databases and normalized to have straightforward and consistent formats, naming, and semantics. The goal of URIEL and lang2vec is to enable multilingual NLP, especially on less-resourced languages and make possible types of experiments (especially but not exclusively related to NLP tasks) that are otherwise difficult or impossible due to the sparsity and incommensurability of the data sources. lang2vec vectors have been shown to reduce perplexity in multilingual language modeling, when compared to one-hot language identification vectors.