Is Part-of-Speech Tagging a Solved Task? An Evaluation of POS Taggers for the German Web as Corpus
Is Part-of-Speech Tagging a Solved Task? An Evaluation of POS Taggers for the German Web as Corpus
复制标题
词性标记是一项已解决的任务吗?
DOI:
--
复制
发表时间:
2009
期刊:
影响因子:
--
通讯作者:
S. Evert
中科院分区:
文献类型:
--
作者:
S. Evert
Part-of-speech (POS) tagging is an important preprocessing step in natural language processing. It is often considered to be a “solved task”, with published tagging accuracies around 97%. Our evaluation of five state-of-the-art POS taggers on German Web texts shows that such high accuracies can only be achieved under artificial cross-validation conditions. In a real-life scenario, accuracy drops below 93% with enormous variation between different text genres, making the taggers unsuitable for fully automatic processing. We find that HMM taggers are more robust and much faster than advanced machine-learning approaches such as MaxEnt. Promising directions for future research are unsupervised learning of a tagger lexicon from large unannotated corpora, as well as developing adaptive tagging models.