Lightweight Cross-Lingual Sentence Representation Learning

Lightweight Cross-Lingual Sentence Representation Learning
复制标题

DOI:
10.18653/v1/2021.acl-long.226
复制
发表时间:
2021-05
期刊:
Malware Analysis Using Artificial Intelligence and Deep Learning
影响因子:
--
通讯作者:
Zhuoyuan Mao;Prakhar Gupta;Chenhui Chu;Martin Jaggi;S. Kurohashi
Zhuoyuan Mao;Prakhar Gupta;Chenhui Chu;Martin Jaggi;S. Kurohashi
中科院分区:
其他
文献类型:
--
作者:
Zhuoyuan Mao;Prakhar Gupta;Chenhui Chu;Martin Jaggi;S. Kurohashi

文献摘要

被引文献

相似文献

用于学习固定维度跨语言句子表示的大型模型(例如 LASER)(Artetxe 和 Schwenk,2019b)可以显着提高下游任务的性能。然而,由于内存限制,基于此类大规模模型的进一步增加和修改通常是不切实际的。在这项工作中,我们引入了一种轻量级的双变压器架构,只有 2 层,用于生成内存高效的跨语言句子表示。我们探索了不同的训练任务,并观察到当前的跨语言训练任务对于这种浅层架构还有很多不足之处。为了改善这一问题,我们提出了一种新颖的跨语言语言模型,它将现有的单词掩码语言模型与新提出的跨语言标记级重建任务相结合。我们通过引入两个计算精简的句子级对比学习任务来进一步增强训练任务,以增强跨语言句子表示空间的对齐,从而弥补生成任务的轻量级变压器的学习瓶颈。我们与跨语言句子检索和多语言文档分类的竞争模型进行比较,证实了新提出的浅层模型训练任务的有效性。
Large-scale models for learning fixed-dimensional cross-lingual sentence representations like LASER (Artetxe and Schwenk, 2019b) lead to significant improvement in performance on downstream tasks. However, further increases and modifications based on such large-scale models are usually impractical due to memory limitations. In this work, we introduce a lightweight dual-transformer architecture with just 2 layers for generating memory-efficient cross-lingual sentence representations. We explore different training tasks and observe that current cross-lingual training tasks leave a lot to be desired for this shallow architecture. To ameliorate this, we propose a novel cross-lingual language model, which combines the existing single-word masked language model with the newly proposed cross-lingual token-level reconstruction task. We further augment the training task by the introduction of two computationally-lite sentence-level contrastive learning tasks to enhance the alignment of cross-lingual sentence representation space, which compensates for the learning bottleneck of the lightweight transformer for generative tasks. Our comparisons with competing models on cross-lingual sentence retrieval and multilingual document classification confirm the effectiveness of the newly proposed training tasks for a shallow model.