Normalization of Language Embeddings for Cross-Lingual Alignment

Normalization of Language Embeddings for Cross-Lingual Alignment
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
P. Aboagye;Yan Zheng;Chin-Chia Michael Yeh;Junpeng Wang;Wei Zhang;Liang Wang;Hao Yang;J. M. Phillips
P. Aboagye;Yan Zheng;Chin-Chia Michael Yeh;Junpeng Wang;Wei Zhang;Liang Wang;Hao Yang;J. M. Phillips
中科院分区:
其他
文献类型:
--
作者:
P. Aboagye;Yan Zheng;Chin-Chia Michael Yeh;Junpeng Wang;Wei Zhang;Liang Wang;Hao Yang;J. M. Phillips

文献摘要

相似文献

学习一个良好的转换函数,将两种语言的词向量映射到一个共享的跨语言词向量空间,在跨语言自然语言处理(NLP)中起着至关重要的作用。它在翻译任务中很有用,并且对于将基于像英语这样的高资源语言构建的复杂模型直接应用于对齐的低资源语言非常重要。虽然普罗克汝斯忒斯(Procrustes)方法和其他技术在对齐语言模型方面取得了一些成功,但最近人们发现结构差异(例如,由于词频不同)为各种单语嵌入创造了不同的特征。当这些特征在不同语言之间存在差异时,它与语言的对齐程度以及它们在跨语言下游任务中的表现相关。在这项工作中,我们开发了一种非常通用的语言嵌入归一化程序,它综合并包含了以前的各种方法,该程序在不破坏语言内在含义的情况下消除了不同语言之间的这些结构特征。我们证明在相似性、翻译和跨语言分类任务中含义得以保留,并且对齐得到了改善。我们提出的归一化方法在每个任务以及每种对齐方法上明显优于所有先前的方法,如中心化和向量归一化。
Learning a good transfer function to map the word vectors from two languages 1 into a shared cross-lingual word vector space plays a crucial role in cross-lingual 2 NLP. It is useful in translation tasks and important in allowing complex models 3 built on a high-resource language like English to be directly applied on an aligned 4 low resource language. While Procrustes and other techniques can align language 5 models with some success, it has recently been identified that structural differences 6 (for instance, due to differing word frequency) create different profiles for various 7 monolingual embedding. When these profiles differ across languages, it corre-8 lates with how well languages can align and their performance on cross-lingual 9 downstream tasks. In this work, we develop a very general language embedding 10 normalization procedure, building and subsuming various previous approaches, 11 which removes these structural profiles across languages without destroying their 12 intrinsic meaning. We demonstrate that meaning is retained and alignment is 13 improved on similarity, translation, and cross-language classification tasks. Our 14 proposed normalization clearly outperforms all prior approaches like centering and 15 vector normalization on each task and with each alignment approach. 16