Making Monolingual Sentence Embeddings Multilingual Using Knowledge Distillation

Making Monolingual Sentence Embeddings Multilingual Using Knowledge Distillation
复制标题

DOI:
10.18653/v1/2020.emnlp-main.365
复制
发表时间:
2020-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Nils Reimers;Iryna Gurevych
Nils Reimers;Iryna Gurevych
中科院分区:
其他
文献类型:
--
作者:
Nils Reimers;Iryna Gurevych

文献摘要

被引文献

相似文献

我们提出了一种简单有效的方法,将现有的句子嵌入模型扩展到新的语言中。这允许从以前的单语言模型创建多语言版本。训练是基于翻译后的句子应该映射到与原始句子在向量空间中的相同位置的想法。我们使用原始(单语)模型为源语言生成句子嵌入,然后在翻译的句子上训练一个新的系统来模仿原始模型。与其他训练多语言句子嵌入的方法相比,这种方法有几个优点:它很容易用相对较少的样本将现有模型扩展到新的语言,更容易确保向量空间所需的属性,并且训练的硬件要求更低。我们对来自不同语系的10种语言演示了我们的方法的有效性。将句子嵌入模型扩展到400多种语言的代码是公开的。
We present an easy and efficient method to extend existing sentence embedding models to new languages. This allows to create multilingual versions from previously monolingual models. The training is based on the idea that a translated sentence should be mapped to the same location in the vector space as the original sentence. We use the original (monolingual) model to generate sentence embeddings for the source language and then train a new system on translated sentences to mimic the original model. Compared to other methods for training multilingual sentence embeddings, this approach has several advantages: It is easy to extend existing models with relatively few samples to new languages, it is easier to ensure desired properties for the vector space, and the hardware requirements for training is lower. We demonstrate the effectiveness of our approach for 10 languages from various language families. Code to extend sentence embeddings models to more than 400 languages is publicly available.