An Empirical Study of the Downstream Reliability of Pre-Trained Word Embeddings
An Empirical Study of the Downstream Reliability of Pre-Trained Word Embeddings
复制标题
DOI:
10.18653/v1/2020.coling-main.299
复制
发表时间:
2020-12
期刊:
影响因子:
--
通讯作者:
Anthony Rios;Brandon Lwowski
中科院分区:
文献类型:
--
作者:
Anthony Rios;Brandon Lwowski
While pre-trained word embeddings have been shown to improve the performance of downstream tasks, many questions remain regarding their reliability: Do the same pre-trained word embeddings result in the best performance with slight changes to the training data? Do the same pre-trained embeddings perform well with multiple neural network architectures? Do imputation strategies for unknown words impact reliability? In this paper, we introduce two new metrics to understand the downstream reliability of word embeddings. We find that downstream reliability of word embeddings depends on multiple factors, including, the evaluation metric, the handling of out-of-vocabulary words, and whether the embeddings are fine-tuned.