Self-paced Multi-view Co-training

Self-paced Multi-view Co-training
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Fan Ma;Deyu Meng;Xuanyi Dong;Yi Yang;S. Kaski
Fan Ma;Deyu Meng;Xuanyi Dong;Yi Yang;S. Kaski
中科院分区:
其他
文献类型:
--
作者:
Fan Ma;Deyu Meng;Xuanyi Dong;Yi Yang;S. Kaski

文献摘要

被引文献

相似文献

协同训练是一种著名的半监督学习方法,它在两个或多个不同的视图上训练分类器,并以迭代的方式交换未标记实例的伪标签。在协同训练过程中,特别是在初始训练阶段,未标记的实例的伪标签很可能是假的,而标准的协同训练算法采用“draw without replacement”的策略,不会将这些错误标记的实例从训练阶段移除。此外,传统的协同训练方法大多是针对双视图场景实现的,其在多视图场景下的扩展并不直观。这些问题不仅降低了它们的性能和应用范围,而且阻碍了它们的基本理论。此外,没有优化模型来解释协同训练过程设法优化的目标。为了解决这些问题,在本研究中,我们设计了一个统一的自定进度多视图协同训练(SPamCo)框架,该框架绘制了未标记的替换实例。提出了两个指定的协正则化项,以制定不同的策略来选择训练过程中的伪标记实例。两种形式具有相同的优化策略,这与协同训练中的迭代过程一致,可以自然地扩展到多视图场景。引入分布式优化策略,并行训练各视图的分类器,进一步提高算法的效率。此外,还证明了SPamCo算法是PAC可学习的,支持了其理论的合理性。在合成、文本分类、人物再识别、图像识别和目标检测数据集上进行的实验证明了该方法的优越性。
Co-training is a well-known semi-supervised learning approach which trains classifiers on two or more different views and exchanges pseudo labels of unlabeled instances in an iterative way. During the co-training process, pseudo labels of unlabeled instances are very likely to be false especially in the initial training, while the standard co-training algorithm adopts a “draw without replacement” strategy and does not remove these wrongly labeled instances from training stages. Besides, most of the traditional co-training approaches are implemented for two-view cases, and their extensions in multi-view scenarios are not intuitive. These issues not only degenerate their performance as well as available application range but also hamper their fundamental theory. Moreover, there is no optimization model to explain the objective a co-training process manages to optimize. To address these issues, in this study we design a unified self-paced multi-view co-training (SPamCo) framework which draws unlabeled instances with replacement. Two specified co-regularization terms are formulated to develop different strategies for selecting pseudo-labeled instances during training. Both forms share the same optimization strategy which is consistent with the iteration process in co-training and can be naturally extended to multi-view scenarios. A distributed optimization strategy is also introduced to train the classifier of each view in parallel to further improve the efficiency of the algorithm. Furthermore, the SPamCo algorithm is proved to be PAC learnable, supporting its theoretical soundness. Experiments conducted on synthetic, text categorization, person re-identification, image recognition and object detection data sets substantiate the superiority of the proposed method.