Bitext Mining Using Distilled Sentence Representations for Low-Resource Languages

Bitext Mining Using Distilled Sentence Representations for Low-Resource Languages
复制标题

使用低资源语言的精炼句子表示进行双文本挖掘

DOI:
10.18653/v1/2022.findings-emnlp.154
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
Holger Schwenk
Holger Schwenk
中科院分区:
--
文献类型:
--
作者:
Kevin Heffernan;Onur cCelebi;Holger Schwenk

文献摘要

被引文献

相似文献

将多语言表示学习扩展到数百种最常用的语言之外是具有挑战性的,特别是要涵盖资源较少的语言的长尾。一种有希望的方法是训练能够跨语言转移的一对一的多语言模式,但这些模式经常受到能力不足和无关语言之间的干扰。相反,我们放弃了这种方法,将重点放在训练多种语言(家族)特定的表示上,但最重要的是使所有语言仍然能够在相同的表示空间中编码。为了实现这一点,我们专注于教师-学生培训,允许所有编码器相互兼容以进行比特文本挖掘,并支持快速学习新语言。我们提出了一种新的教师-学生训练方案,将监督训练和自我监督训练相结合,使编码者能够利用单语训练数据,这在资源不足的情况下是有价值的。我们的方法明显优于原始的激光编码器。我们研究资源非常少的语言,处理50种非洲语言,其中许多语言没有被任何其他模型涵盖。对于这些语言,我们训练句子编码器,挖掘比特文本,并通过训练NMT系统来验证比特文本。
Scaling multilingual representation learning beyond the hundred most frequent languages is challenging, in particular to cover the long tail of low-resource languages. A promising approach has been to train one-for-all multilingual models capable of cross-lingual transfer, but these models often suffer from insufficient capacity and interference between unrelated languages. Instead, we move away from this approach and focus on training multiple language (family) specific representations, but most prominently enable all languages to still be encoded in the same representational space. To achieve this, we focus on teacher-student training, allowing all encoders to be mutually compatible for bitext mining, and enabling fast learning of new languages. We introduce a new teacher-student training scheme which combines supervised and self-supervised training, allowing encoders to take advantage of monolingual training data, which is valuable in the low-resource setting. Our approach significantly outperforms the original LASER encoder. We study very low-resource languages and handle 50 African languages, many of which are not covered by any other model. For these languages, we train sentence encoders, mine bitexts, and validate the bitexts by training NMT systems.