Multilingual Part-of-Speech Tagging with Bidirectional Long Short-Term Memory Models and Auxiliary Loss

Multilingual Part-of-Speech Tagging with Bidirectional Long Short-Term Memory Models and Auxiliary Loss
复制标题

DOI:
10.18653/v1/p16-2067
复制
发表时间:
2016-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Barbara Plank;Anders Søgaard;Yoav Goldberg
Barbara Plank;Anders Søgaard;Yoav Goldberg
中科院分区:
其他
文献类型:
--
作者:
Barbara Plank;Anders Søgaard;Yoav Goldberg

文献摘要

被引文献

相似文献

双向长短期记忆(bi - LSTM)网络最近在各种自然语言处理(NLP)序列建模任务中被证明是成功的,但对于它们对输入表示、目标语言、数据集大小和标签噪声的依赖情况却知之甚少。我们探讨了这些问题,并针对词性标注评估了具有单词、字符和Unicode字节嵌入的双向长短期记忆网络。我们在不同语言和数据规模下将双向长短期记忆网络与传统的词性标注器进行了比较。我们还提出了一种新的双向长短期记忆模型,该模型将词性标注损失函数与一个考虑罕见词的辅助损失函数相结合。该模型在22种语言上取得了最先进的性能,并且对于形态复杂的语言效果尤其好。我们的分析表明,双向长短期记忆网络对训练数据大小和标签损坏(在较小噪声水平下)的敏感度比之前所认为的要低。
Bidirectional long short-term memory (bi-LSTM) networks have recently proven successful for various NLP sequence modeling tasks, but little is known about their reliance to input representations, target languages, data set size, and label noise. We address these issues and evaluate bi-LSTMs with word, character, and unicode byte embeddings for POS tagging. We compare bi-LSTMs to traditional POS taggers across languages and data sizes. We also present a novel bi-LSTM model, which combines the POS tagging loss function with an auxiliary loss function that accounts for rare words. The model obtains state-of-the-art performance across 22 languages, and works especially well for morphologically complex languages. Our analysis suggests that bi-LSTMs are less sensitive to training data size and label corruptions (at small noise levels) than previously assumed.