Leveraging Pre-Trained Embeddings for Welsh Taggers

Leveraging Pre-Trained Embeddings for Welsh Taggers
复制标题

DOI:
10.18653/v1/w19-4332
复制
发表时间:
2019-08
期刊:
--
影响因子:
--
通讯作者:
I. Ezeani;S. Piao;Steven Neale;Paul Rayson;Dawn Knight
I. Ezeani;S. Piao;Steven Neale;Paul Rayson;Dawn Knight
中科院分区:
其他
文献类型:
--
作者:
I. Ezeani;S. Piao;Steven Neale;Paul Rayson;Dawn Knight

文献摘要

被引文献

相似文献

虽然单词嵌入模型在下游自然语言处理(NLP)任务中的应用已被证明是成功的,但由于缺乏足够的数据来训练模型,低资源语言的好处在一定程度上受到限制。然而,针对低资源语言的NLP研究工作一直专注于不断寻求利用预训练模型的方法,以提高为处理这些语言而构建的NLP系统的性能,而无需重新发明轮子。其中一种语言是威尔士语,因此,在本文中,我们提出了我们的实验结果,使用FastText的预训练嵌入模型学习一个简单的多任务神经网络模型,用于威尔士语的词性和语义标记。我们的模型的性能进行了比较,与现有的基于规则的独立标记部分的语音和语义标记。尽管它的简单性和能力,同时执行这两项任务,我们的tagger相比,非常好与现有的taggers。
While the application of word embedding models to downstream Natural Language Processing (NLP) tasks has been shown to be successful, the benefits for low-resource languages is somewhat limited due to lack of adequate data for training the models. However, NLP research efforts for low-resource languages have focused on constantly seeking ways to harness pre-trained models to improve the performance of NLP systems built to process these languages without the need to re-invent the wheel. One such language is Welsh and therefore, in this paper, we present the results of our experiments on learning a simple multi-task neural network model for part-of-speech and semantic tagging for Welsh using a pre-trained embedding model from FastText. Our model’s performance was compared with those of the existing rule-based stand-alone taggers for part-of-speech and semantic taggers. Despite its simplicity and capacity to perform both tasks simultaneously, our tagger compared very well with the existing taggers.