Subword-Based Compact Reconstruction for Open-Vocabulary Neural Word Embeddings

Subword-Based Compact Reconstruction for Open-Vocabulary Neural Word Embeddings
复制标题

开放词汇神经词嵌入的基于子词的紧凑重建

DOI:
10.1109/taslp.2021.3125133
复制
发表时间:
2021
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Inui Kentaro
Inui Kentaro
中科院分区:
--
文献类型:
--
作者:
Sasaki Shota;Suzuki Jun;Inui Kentaro

文献摘要

相似文献

神经词嵌入方法已成为解决人工智能研究领域许多应用问题的重要基础资源。它们已被证明能够在向量空间中捕获高质量的句法和语义关系。尽管神经词嵌入具有重要的影响,但它也有一些缺点。在本文中,我们主要关注关于训练良好的词嵌入的两个问题:(i)大量的内存需求和(ii)词汇外(OOV)词的不适用性。为了克服这两个问题,我们提出了一种通过使用子词信息重构预训练词嵌入的方法,该方法在相当小的固定空间内有效地表示大量子词嵌入,同时防止原始词嵌入的质量下降。该方法的关键技术有两个:内存共享嵌入和键值查询自关注机制的变体。我们的实验表明,我们重建的基于子词的词嵌入可以在一个小的固定空间内成功地模仿训练良好的词嵌入,同时防止多个语言基准数据集的质量下降,并可以同时预测OOV词的有效嵌入。我们还证明了我们的重建方法在应用于下游任务(如命名实体识别和自然语言推理任务)时的有效性。
The methodology of neural word embeddings has become an important fundamental resource for tackling many applications in the artificial intelligence (AI) research field. They have successfully been proven to capture high-quality syntactic and semantic relationships in a vector space. Despite their significant impact, neural word embeddings have several disadvantages. In this paper, we focus on two issues regarding well-trained word embeddings: (i) the massive memory requirement and (ii) the inapplicability of out-of-vocabulary (OOV) words. To overcome these two issues, we propose a method of reconstructing pre-trained word embeddings by using subword information that effectively represents a large number of subword embeddings in a considerably small fixed space while preventing quality degradation from the original word embeddings. The key techniques of our method are twofold: memory-shared embeddings and a variant of the key-value-query self-attention mechanism. Our experiments show that our reconstructed subword-based word embeddings can successfully imitate well-trained word embeddings in a small fixed space while preventing quality degradation across several linguistic benchmark datasets and can simultaneously predict effective embeddings of OOV words. We also demonstrate the effectiveness of our reconstruction method when it is applied to downstream tasks, such as named entity recognition and natural language inference tasks.