Semi-supervised learning integrated with classifier combination for word sense disambiguation

Semi-supervised learning integrated with classifier combination for word sense disambiguation
复制标题

DOI:
10.1016/j.csl.2007.11.001
复制
发表时间:
2008-10-01
影响因子:
4.3
通讯作者:
Nguyen, Le-Minh
Nguyen, Le-Minh
中科院分区:
计算机科学3区
文献类型:
--
作者:
Le, Anh-Cuong;Shimazu, Akira;Nguyen, Le-Minh

文献摘要

被引文献

相似文献

词义消歧是指确定一个多义词在特定语境中的正确词义的问题。本文研究了在半监督学习的框架内使用未标记数据进行词义消歧,其中标记数据是从未标记数据迭代扩展的。针对这种方法,我们首先明确识别和分析了一般自举算法中固有的三个问题;即训练数据的不平衡、新标记示例的置信度以及最终分类器的生成;所有这些都将在一个共同的自举框架内综合考虑。然后,我们提出了解决这些问题的分类器组合策略的帮助下。这导致了一般自举算法的几个新的变体。在英语Senseval-2和Senseval-3词汇样本上进行的实验表明,与以往的研究相比,所提出的解决方案是有效的,并显着提高监督式词义消歧。(c)2007爱思唯尔有限公司保留所有权利。
Word sense disambiguation (WSD) is the problem of determining the right sense of a polysernous word in a certain context. This paper investigates the use of unlabeled data for WSD within a framework of semi-supervised learning, in which labeled data is iteratively extended from unlabeled data. Focusing on this approach, we first explicitly identify and analyze three problems inherently occurred piecemeal in the general bootstrapping algorithm; namely the imbalance of training data, the confidence of new labeled examples, and the final classifier generation; all of which will be considered integratedly within a common framework of bootstrapping. We then propose solutions for these problems with the help of classifier combination strategies. This results in several new variants of the general bootstrapping algorithm. Experiments conducted on the English lexical samples of Senseval-2 and Senseval-3 show that the proposed solutions are effective in comparison with previous studies, and significantly improve supervised WSD. (c) 2007 Elsevier Ltd. All rights reserved.