Normalization of Language Embeddings for Cross-Lingual Alignment
Normalization of Language Embeddings for Cross-Lingual Alignment
复制标题
DOI:
--
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
P. Aboagye;Yan Zheng;Chin-Chia Michael Yeh;Junpeng Wang;Wei Zhang;Liang Wang;Hao Yang;J. M. Phillips
中科院分区:
文献类型:
--
作者:
P. Aboagye;Yan Zheng;Chin-Chia Michael Yeh;Junpeng Wang;Wei Zhang;Liang Wang;Hao Yang;J. M. Phillips
Learning a good transfer function to map the word vectors from two languages 1 into a shared cross-lingual word vector space plays a crucial role in cross-lingual 2 NLP. It is useful in translation tasks and important in allowing complex models 3 built on a high-resource language like English to be directly applied on an aligned 4 low resource language. While Procrustes and other techniques can align language 5 models with some success, it has recently been identified that structural differences 6 (for instance, due to differing word frequency) create different profiles for various 7 monolingual embedding. When these profiles differ across languages, it corre-8 lates with how well languages can align and their performance on cross-lingual 9 downstream tasks. In this work, we develop a very general language embedding 10 normalization procedure, building and subsuming various previous approaches, 11 which removes these structural profiles across languages without destroying their 12 intrinsic meaning. We demonstrate that meaning is retained and alignment is 13 improved on similarity, translation, and cross-language classification tasks. Our 14 proposed normalization clearly outperforms all prior approaches like centering and 15 vector normalization on each task and with each alignment approach. 16