Attentive Mimicking: Better Word Embeddings by Attending to Informative Contexts

Attentive Mimicking: Better Word Embeddings by Attending to Informative Contexts
复制标题

细心模仿:通过关注信息丰富的上下文来获得更好的词嵌入

DOI:
--
复制
发表时间:
2019
期刊:
North American Chapter of the Association for Computational Linguistics
影响因子:
--
通讯作者:
Hinrich Schütze
Hinrich Schütze
中科院分区:
--
文献类型:
--
作者:
Timo Schick;Hinrich Schütze

文献摘要

被引文献

相似文献

由于上下文信息的稀疏,学习高质量的稀有词嵌入是一个困难的问题。模仿(Pinter等人,2017年)已经提出了一种解决方案:给定通过标准算法学习的嵌入,首先训练模型以从其表面形式再现频繁单词的嵌入,然后用于计算罕见单词的嵌入。在本文中,我们介绍了注意模仿:模仿模型不仅可以访问一个词的表面形式,而且可以访问所有可用的上下文,并学习参加最翔实和可靠的上下文计算嵌入。在四个任务的评估,我们表明,注意模仿优于以前的工作,罕见的和中频的话。因此,与以前的工作相比,专注模仿改善了更大部分词汇的嵌入,包括中频范围。
Learning high-quality embeddings for rare words is a hard problem because of sparse context information. Mimicking (Pinter et al., 2017) has been proposed as a solution: given embeddings learned by a standard algorithm, a model is first trained to reproduce embeddings of frequent words from their surface form and then used to compute embeddings for rare words. In this paper, we introduce attentive mimicking: the mimicking model is given access not only to a word’s surface form, but also to all available contexts and learns to attend to the most informative and reliable contexts for computing an embedding. In an evaluation on four tasks, we show that attentive mimicking outperforms previous work for both rare and medium-frequency words. Thus, compared to previous work, attentive mimicking improves embeddings for a much larger part of the vocabulary, including the medium-frequency range.