Creating Welsh Language Word Embeddings

Creating Welsh Language Word Embeddings
复制标题

DOI:
10.3390/app11156896
复制
发表时间:
2021-07
期刊:
影响因子:
--
通讯作者:
P. Corcoran;Geraint I. Palmer;Laura Arman;Dawn Knight;Irena Spasic
P. Corcoran;Geraint I. Palmer;Laura Arman;Dawn Knight;Irena Spasic
中科院分区:
--
文献类型:
--
作者:
P. Corcoran;Geraint I. Palmer;Laura Arman;Dawn Knight;Irena Spasic

文献摘要

被引文献

相似文献

词嵌入是词在向量空间中的表示,该向量空间通过距离和方向对词之间的语义关系进行建模。在这项研究中,我们采用了两种现有的方法,word2vec和fastText,自动学习威尔士语词嵌入考虑到这种语言的句法和形态的特质。这些方法利用分布式语义的原则,因此,需要一个大的语料库进行训练。然而,威尔士语是一种少数民族语言,因此与英语相比,威尔士语数据明显较少。因此,组装一个足够大的文本语料库不是一个简单的努力。尽管如此,我们还是从11个来源中汇编了92,963,671个单词的语料库,这是威尔士语最大的语料库。威尔士语标点符号的相对复杂性使得该语料库的标记化相对具有挑战性,因为标点符号不能用于边界检测。我们考虑了几种标记化方法,包括一种专门为威尔士语设计的方法。为了解释丰富的变形,我们使用了一种基于子词的词嵌入学习方法,因此可以在训练阶段更有效地将不同的表面形式联系起来。我们进行了定性和定量的评价所产生的词嵌入,这优于以前描述的词嵌入在威尔士语的一部分,包括157种语言的更大的研究。我们的研究是第一个专门关注威尔士语单词嵌入的研究。
Word embeddings are representations of words in a vector space that models semantic relationships between words by means of distance and direction. In this study, we adapted two existing methods, word2vec and fastText, to automatically learn Welsh word embeddings taking into account syntactic and morphological idiosyncrasies of this language. These methods exploit the principles of distributional semantics and, therefore, require a large corpus to be trained on. However, Welsh is a minoritised language, hence significantly less Welsh language data are publicly available in comparison to English. Consequently, assembling a sufficiently large text corpus is not a straightforward endeavour. Nonetheless, we compiled a corpus of 92,963,671 words from 11 sources, which represents the largest corpus of Welsh. The relative complexity of Welsh punctuation made the tokenisation of this corpus relatively challenging as punctuation could not be used for boundary detection. We considered several tokenisation methods including one designed specifically for Welsh. To account for rich inflection, we used a method for learning word embeddings that is based on subwords and, therefore, can more effectively relate different surface forms during the training phase. We conducted both qualitative and quantitative evaluation of the resulting word embeddings, which outperformed previously described word embeddings in Welsh as part of larger study including 157 languages. Our study was the first to focus specifically on Welsh word embeddings.