Tagging with Small Training Corpora

Tagging with Small Training Corpora
复制标题

使用小型训练语料库进行标记

DOI:
10.1007/3-540-44816-0_7
复制
发表时间:
2001
期刊:
International Symposium on Intelligent Data Analysis
影响因子:
--
通讯作者:
J. Lopes
J. Lopes
中科院分区:
--
文献类型:
--
作者:
N. Marques;J. Lopes

文献摘要

被引文献

相似文献

文本数据的分析可以从使用预定义的标记集对单词进行分类开始。然而,对于自然语言文本来说,如何在不受限制的文本中为单词分配词性标签(称为POS-tagging)仍然是一个问题。目前大部分标注器需要大量的手工标注文本进行训练(约105个预标注单词):这需要语言上受过高度训练的人力来完成一项高度重复和枯燥的工作,并且所获得的结果没有最佳质量。此外,当一个人想要转换到另一种文本类型时,同样的问题必须再次面对。我们的建议是相反的。通过将大型词典与高效的基于神经网络的标注器生成器相结合,我们可以生成post -标注器,使用不超过104个手动校正的标注词进行训练。这种训练标记的文本大小可以手工修正。给出了SUSANNE语料库的实验结果并进行了讨论。另外三种不同的葡萄牙语语料库的结果也进行了讨论。当测试集中出现未知词时,准确率达到96%。当测试集中的每个单词都已知时,准确率达到98%。
The analysis of textual data may start by classifying words usinga predefined tag set. However, it is still a problem for natural language text understanding the assignment of part-of-speech tags to words in unrestricted text (called POS-tagging). Most part of current taggers require huge amounts of hand tagged text for training (in the order of 105pretagged words): it requires linguistically highly trained man power for a highly repetitive and boring job, and the results obtained have no optimal quality. Moreover, when one wants to change to another text genre the same kind of problem must be faced again. Our proposal goes in another direction. By carefully combininga large lexicon with an efficient neural network based generator of taggers we can generate POS-taggers using no more than 104hand corrected tagged words for training. This training tagged text size can be feasibly hand corrected. Experimental results are presented and discussed for the SUSANNE Corpus. Results in three additional different Portuguese corpora are also discussed. 96% precision rates are obtained when unknown words occur in the test set. 98% precision rates are obtained when every word in the test set is known.