Is Part-of-Speech Tagging a Solved Task? An Evaluation of POS Taggers for the German Web as Corpus

Is Part-of-Speech Tagging a Solved Task? An Evaluation of POS Taggers for the German Web as Corpus
复制标题

词性标记是一项已解决的任务吗?

DOI:
--
复制
发表时间:
2009
期刊:
影响因子:
--
通讯作者:
S. Evert
S. Evert
中科院分区:
--
文献类型:
--
作者:
S. Evert

文献摘要

被引文献

相似文献

词性标注是自然语言处理中一个重要的预处理步骤。这通常被认为是一项“已解决的任务”,发布的标注准确率约为97%。我们在德国Web文本上对五个最先进的POS标记器的评估表明,只有在人工交叉验证条件下才能达到如此高的准确率。在现实生活中,由于不同文本体裁之间的巨大差异,准确率降至93%以下,使得标记器不适合全自动处理。我们发现,HMM标记器比MaxEnt等先进的机器学习方法更健壮,速度也更快。未来研究的有希望的方向是从大型未注释语料库中无监督地学习标记器词典,以及开发自适应标注模型。
Part-of-speech (POS) tagging is an important preprocessing step in natural language processing. It is often considered to be a “solved task”, with published tagging accuracies around 97%. Our evaluation of five state-of-the-art POS taggers on German Web texts shows that such high accuracies can only be achieved under artificial cross-validation conditions. In a real-life scenario, accuracy drops below 93% with enormous variation between different text genres, making the taggers unsuitable for fully automatic processing. We find that HMM taggers are more robust and much faster than advanced machine-learning approaches such as MaxEnt. Promising directions for future research are unsupervised learning of a tagger lexicon from large unannotated corpora, as well as developing adaptive tagging models.