Semi-Supervised Sequential Labeling and Segmentation Using Giga-Word Scale Unlabeled Data

Semi-Supervised Sequential Labeling and Segmentation Using Giga-Word Scale Unlabeled Data
复制标题

DOI:
--
复制
发表时间:
2008-06
期刊:
--
影响因子:
--
通讯作者:
Jun Suzuki;Hideki Isozaki
Jun Suzuki;Hideki Isozaki
中科院分区:
其他
文献类型:
--
作者:
Jun Suzuki;Hideki Isozaki

文献摘要

被引文献

相似文献

本文提供的证据表明,在半监督学习中使用更多的未标记数据可以提高自然语言处理(NLP)任务的性能,例如词性标记,句法组块和命名实体识别。我们首先提出了一个简单而强大的半监督判别模型,适合处理大规模的未标记数据。然后,我们描述了广泛使用的测试集合,即PTB III数据,CoNLL'00和'03共享任务数据的上述三个NLP任务,分别进行实验。我们整合了多达1G字(10亿个标记)的未标记数据,这是有史以来用于这些任务的最大数量的未标记数据,以研究性能改进。此外,我们的结果是上级的最佳报告的结果,所有上述测试集合。
This paper provides evidence that the use of more unlabeled data in semi-supervised learning can improve the performance of Natural Language Processing (NLP) tasks, such as part-of-speech tagging, syntactic chunking, and named entity recognition. We first propose a simple yet powerful semi-supervised discriminative model appropriate for handling large scale unlabeled data. Then, we describe experiments performed on widely used test collections, namely, PTB III data, CoNLL’00 and ’03 shared task data for the above three NLP tasks, respectively. We incorporate up to 1G-words (one billion tokens) of unlabeled data, which is the largest amount of unlabeled data ever used for these tasks, to investigate the performance improvement. In addition, our results are superior to the best reported results for all of the above test collections.