CoTrade: Confident Co-Training With Data Editing

CoTrade: Confident Co-Training With Data Editing
复制标题

DOI:
10.1109/tsmcb.2011.2157998
复制
发表时间:
2011-12
期刊:
IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics)
影响因子:
--
通讯作者:
Min-Ling Zhang;Zhi-Hua Zhou
Min-Ling Zhang;Zhi-Hua Zhou
中科院分区:
其他
文献类型:
--
作者:
Min-Ling Zhang;Zhi-Hua Zhou

文献摘要

被引文献

相似文献

协同训练是主要的半监督学习范例之一,它在两个不同的视图上迭代训练两个分类器,并使用任一分类器对未标记示例的预测来增强另一个分类器的训练集。在协同训练过程中,特别是在分类器的准确度一般的第一轮中,一个分类器很可能会收到另一个分类器错误预测的未标记示例的标签。因此,协同训练型算法的性能通常不稳定。在本文中,如何在不同视图之间可靠地传递标签信息的问题通过一种名为 COTRADE 的新型协同训练算法来解决。在每一轮标签中,COTRADE 分两步执行标签通信过程。首先,基于特定的数据编辑技术明确估计任一分类器对未标记示例的预测的置信度。其次,任一分类器具有较高置信度的多个预测标签被传递到另一个分类器,其中施加某些约束以避免引入不需要的分类噪声。对三个领域的多个真实数据集的实验表明,COTRADE 可以有效地利用未标记的数据来实现更好的泛化性能。
Co-training is one of the major semi-supervised learning paradigms that iteratively trains two classifiers on two different views, and uses the predictions of either classifier on the unlabeled examples to augment the training set of the other. During the co-training process, especially in initial rounds when the classifiers have only mediocre accuracy, it is quite possible that one classifier will receive labels on unlabeled examples erroneously predicted by the other classifier. Therefore, the performance of co-training style algorithms is usually unstable. In this paper, the problem of how to reliably communicate labeling information between different views is addressed by a novel co-training algorithm named COTRADE. In each labeling round, COTRADE carries out the label communication process in two steps. First, confidence of either classifier's predictions on unlabeled examples is explicitly estimated based on specific data editing techniques. Secondly, a number of predicted labels with higher confidence of either classifier are passed to the other one, where certain constraints are imposed to avoid introducing undesirable classification noise. Experiments on several real-world datasets across three domains show that COTRADE can effectively exploit unlabeled data to achieve better generalization performance.