Tagging with Small Training Corpora
Tagging with Small Training Corpora
复制标题
使用小型训练语料库进行标记
DOI:
10.1007/3-540-44816-0_7
复制
发表时间:
2001
期刊:
影响因子:
--
通讯作者:
J. Lopes
中科院分区:
文献类型:
--
作者:
N. Marques;J. Lopes
The analysis of textual data may start by classifying words usinga predefined tag set. However, it is still a problem for natural language text understanding the assignment of part-of-speech tags to words in unrestricted text (called POS-tagging). Most part of current taggers require huge amounts of hand tagged text for training (in the order of 105pretagged words): it requires linguistically highly trained man power for a highly repetitive and boring job, and the results obtained have no optimal quality. Moreover, when one wants to change to another text genre the same kind of problem must be faced again. Our proposal goes in another direction. By carefully combininga large lexicon with an efficient neural network based generator of taggers we can generate POS-taggers using no more than 104hand corrected tagged words for training. This training tagged text size can be feasibly hand corrected. Experimental results are presented and discussed for the SUSANNE Corpus. Results in three additional different Portuguese corpora are also discussed. 96% precision rates are obtained when unknown words occur in the test set. 98% precision rates are obtained when every word in the test set is known.