Evaluating Word Embeddings in Extremely Under-Resourced Languages: A Case Study in Bribri

Evaluating Word Embeddings in Extremely Under-Resourced Languages: A Case Study in Bribri
复制标题

评估资源极度贫乏语言中的词嵌入:Bribri 的案例研究

DOI:
--
复制
发表时间:
2022
期刊:
International Conference on Computational Linguistics
影响因子:
--
通讯作者:
Rolando Coto
Rolando Coto
中科院分区:
--
文献类型:
--
作者:
Rolando Coto

文献摘要

被引文献

相似文献

词嵌入对于许多 NLP 任务至关重要,但它们在实际资源贫乏环境中的评估需要进一步检查。本文介绍了 Bribri(一种来自哥斯达黎加的 Chibchan 语言)的案例研究。四个实验改编自英语:单词相似度、WordSim353 相关性、单数任务和类比。在这里,我们讨论它们对资源贫乏的土著语言的适应,并用它们来衡量语义和词法学习。我们使用不同的超参数组合训练了 96 个 word2vec 模型。这种资源贫乏场景的最佳模型是具有中等大小(100 维)和大窗口大小(10)的 Skip-gram。它们与 WordSim353 的平均相关性为 r=0.28,语义奇一的准确率为 76%,结构/形态奇一的准确率为 70%。类比的性能较低:最好的模型可以在大约 60% 的时间内在前 25 个结果中找到适当的语义目标,但只能在 11% 的时间内找到形态/结构目标。未来的研究需要进一步探索形态/结构学习的模式,检查深度学习嵌入的行为,并建立人类基线。该项目旨在改进 Bribri NLP 并最终帮助其维护和振兴。
Word embeddings are critical for numerous NLP tasks but their evaluation in actual under-resourced settings needs further examination. This paper presents a case study in Bribri, a Chibchan language from Costa Rica. Four experiments were adapted from English: Word similarities, WordSim353 correlations, odd-one-out tasks and analogies. Here we discuss their adaptation to an under-resourced Indigenous language and we use them to measure semantic and morphological learning. We trained 96 word2vec models with different hyperparameter combinations. The best models for this under-resourced scenario were Skip-grams with an intermediate size (100 dimensions) and large window sizes (10). These had an average correlation of r=0.28 with WordSim353, a 76% accuracy in semantic odd-one-out and 70% accuracy in structural/morphological odd-one-out. The performance was lower for the analogies: The best models could find the appropriate semantic target amongst the first 25 results approximately 60% of the times, but could only find the morphological/structural target 11% of the times. Future research needs to further explore the patterns of morphological/structural learning, to examine the behavior of deep learning embeddings, and to establish a human baseline. This project seeks to improve Bribri NLP and ultimately help in its maintenance and revitalization.