High Quality ELMo Embeddings for Seven Less-Resourced Languages

High Quality ELMo Embeddings for Seven Less-Resourced Languages
复制标题

适用于七种资源较少的语言的高质量 ELMo 嵌入

DOI:
--
复制
发表时间:
2019
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
M. Robnik
M. Robnik
中科院分区:
--
文献类型:
--
作者:
Matej Ulvcar;M. Robnik

文献摘要

被引文献

相似文献

最近的研究结果表明,在大多数文本分类任务中,使用上下文嵌入的深度神经网络的性能明显优于非上下文嵌入。我们为七种语言提供来自流行的上下文Elmo模型的预计算嵌入:克罗地亚语、爱沙尼亚语、芬兰语、拉脱维亚语、立陶宛语、斯洛文尼亚语和瑞典语。我们证明了嵌入的质量强烈地依赖于训练集的大小,并证明了现有的公开可用的列出语言的ELMO嵌入应该得到改进。我们在更大的训练集上训练新的ELMO嵌入,并展示了它们相对于基线非上下文FastText嵌入的优势。在评估中,我们使用了两个基准,类比任务和NER任务。
Recent results show that deep neural networks using contextual embeddings significantly outperform non-contextual embeddings on a majority of text classification task. We offer precomputed embeddings from popular contextual ELMo model for seven languages: Croatian, Estonian, Finnish, Latvian, Lithuanian, Slovenian, and Swedish. We demonstrate that the quality of embeddings strongly depends on the size of training set and show that existing publicly available ELMo embeddings for listed languages shall be improved. We train new ELMo embeddings on much larger training sets and show their advantage over baseline non-contextual FastText embeddings. In evaluation, we use two benchmarks, the analogy task and the NER task.