Specializing Word Embeddings (for Parsing) by Information Bottleneck

Specializing Word Embeddings (for Parsing) by Information Bottleneck
复制标题

DOI:
10.18653/v1/d19-1276
复制
发表时间:
2019-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Xiang Lisa Li;Jason Eisner
Xiang Lisa Li;Jason Eisner
中科院分区:
其他
文献类型:
--
作者:
Xiang Lisa Li;Jason Eisner

文献摘要

被引文献

相似文献

像埃尔莫和BERT这样的预训练词嵌入包含丰富的语法和语义信息,从而在各种任务上实现最先进的性能。我们提出了一个非常快速的变分信息瓶颈(VIB)方法来非线性压缩这些嵌入,只保留有助于判别式解析器的信息。我们将每个单词嵌入压缩为离散标签或连续向量。在离散的版本中,我们的自动压缩标签形成一个替代的标签集:我们的实验表明,我们的标签捕获传统的POS标签注释中的大部分信息,但我们的标签序列可以更准确地解析在同一级别的标签粒度。在连续的版本中,我们的实验表明,适度压缩的词嵌入我们的方法产生一个更准确的分析器在8 9种语言,不像简单的降维。
Pre-trained word embeddings like ELMo and BERT contain rich syntactic and semantic information, resulting in state-of-the-art performance on various tasks. We propose a very fast variational information bottleneck (VIB) method to nonlinearly compress these embeddings, keeping only the information that helps a discriminative parser. We compress each word embedding to either a discrete tag or a continuous vector. In the discrete version, our automatically compressed tags form an alternative tag set: we show experimentally that our tags capture most of the information in traditional POS tag annotations, but our tag sequences can be parsed more accurately at the same level of tag granularity. In the continuous version, we show experimentally that moderately compressing the word embeddings by our method yields a more accurate parser in 8 of 9 languages, unlike simple dimensionality reduction.