Tri-training: exploiting unlabeled data using three classifiers

Tri-training: exploiting unlabeled data using three classifiers
复制标题

DOI:
10.1109/tkde.2005.186
复制
发表时间:
2005-11
影响因子:
8.9
通讯作者:
Zhi-Hua Zhou;Ming Li
Zhi-Hua Zhou;Ming Li
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zhi-Hua Zhou;Ming Li

文献摘要

被引文献

相似文献

在许多实际的数据挖掘应用程序中,例如Web页面分类,未标记的训练示例很容易获得,但标记的训练示例的获取成本相当高。因此,协同训练等半监督学习算法备受关注。本文提出了一种新的协同训练式半监督学习算法——三训练算法。该算法从原始标记样例集生成三个分类器。然后在三训练过程中使用未标记的示例对这些分类器进行改进。具体来说,在每一轮的三训练中,如果在一定条件下其他两个分类器的标记一致,则为一个分类器标记一个未标记的示例。由于三训练既不需要用足够冗余的视图来描述实例空间,也没有对监督学习算法施加任何约束,因此它的适用性比以前的协同训练风格的算法更广泛。在UCI数据集上的实验和在网页分类任务中的应用表明,三次训练可以有效地利用未标记数据来提高学习性能。
In many practical data mining applications, such as Web page classification, unlabeled training examples are readily available, but labeled ones are fairly expensive to obtain. Therefore, semi-supervised learning algorithms such as co-training have attracted much attention. In this paper, a new co-training style semi-supervised learning algorithm, named tri-training, is proposed. This algorithm generates three classifiers from the original labeled example set. These classifiers are then refined using unlabeled examples in the tri-training process. In detail, in each round of tri-training, an unlabeled example is labeled for a classifier if the other two classifiers agree on the labeling, under certain conditions. Since tri-training neither requires the instance space to be described with sufficient and redundant views nor does it put any constraints on the supervised learning algorithm, its applicability is broader than that of previous co-training style algorithms. Experiments on UCI data sets and application to the Web page classification task indicate that tri-training can effectively exploit unlabeled data to enhance the learning performance.