Combination of Active Learning and Semi-Supervised Learning under a Self-Training Scheme

Combination of Active Learning and Semi-Supervised Learning under a Self-Training Scheme
复制标题

在自我训练方案下的积极学习和半监督学习的结合

DOI:
10.3390/e21100988
复制
发表时间:
2019-10-10
期刊:
影响因子:
2.7
通讯作者:
Kotsiantis S
Kotsiantis S
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Fazakis N;Kanas VG;Aridas CK;Karlos S;Kotsiantis S

文献摘要

参考文献

被引文献

相似文献

影响分类算法性能的主要方面之一是在训练阶段可用的标记数据的量。人们普遍认为,大量数据的标记过程既昂贵又耗时,因为它需要雇用人类的专业知识。对于各种各样的科学领域,未标记的例子很容易收集,但很难以有用的方式处理,从而提高了主题数据集所包含的信息。在这种背景下,各种学习方法已经在文献中研究,旨在有效地利用大量的未标记的数据在学习过程中。最常见的方法通过单独应用主动学习或半监督学习方法来解决这类问题。在这项工作中,主动学习和半监督学习方法的组合,提出了一个共同的自我训练计划下,为了有效地利用可用的未标记的数据。使用有效且鲁棒的熵度量和未标记集的概率分布,以选择最充分的未标记样本用于初始标记集的扩充。通过在55个基准数据集上与监督、半监督和主动学习的基本方法进行比较,验证了该方案的优越性。
One of the major aspects affecting the performance of the classification algorithms is the amount of labeled data which is available during the training phase. It is widely accepted that the labeling procedure of vast amounts of data is both expensive and time-consuming since it requires the employment of human expertise. For a wide variety of scientific fields, unlabeled examples are easy to collect but hard to handle in a useful manner, thus improving the contained information for a subject dataset. In this context, a variety of learning methods have been studied in the literature aiming to efficiently utilize the vast amounts of unlabeled data during the learning process. The most common approaches tackle problems of this kind by individually applying active learning or semi-supervised learning methods. In this work, a combination of active learning and semi-supervised learning methods is proposed, under a common self-training scheme, in order to efficiently utilize the available unlabeled data. The effective and robust metrics of the entropy and the distribution of probabilities of the unlabeled set, to select the most sufficient unlabeled examples for the augmentation of the initial labeled set, are used. The superiority of the proposed scheme is validated by comparing it against the base approaches of supervised, semi-supervised, and active learning in the wide range of fifty-five benchmark datasets.
DOI: 10.1016/j.patrec.2019.07.022
发表时间: 2019-07-01
影响因子: 5.1
作者:
Fazakis, Nikos;Karlos, Stamatis;Sgarbas, Kyriakos
通讯作者: Sgarbas, Kyriakos
DOI: 10.1109/tit.2018.2879897
发表时间: 2019-04-01
影响因子: 2.5
作者:
Anis, Aamir;El Gamal, Aly;Ortega, Antonio
通讯作者: Ortega, Antonio
DOI: 10.1142/s0218213015500335
发表时间: 2016-04-01
影响因子: 1.1
作者:
Amini, Mohammad;Rezaeenour, Jalal;Hadavandi, Esmaeil
通讯作者: Hadavandi, Esmaeil
DOI: 10.1016/s0167-8655(03)00008-4
发表时间: 2003-08-01
影响因子: 5.1
作者:
Chen, YS;Wang, GP;Dong, SH
通讯作者: Dong, SH
DOI: 10.1007/bf00994018
发表时间: 1995-09-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
CORTES, C;VAPNIK, V
通讯作者: VAPNIK, V