Active semi-supervised learning for biological data classification

Active semi-supervised learning for biological data classification
复制标题

DOI:
10.1371/journal.pone.0237428
复制
发表时间:
2020-08-19
期刊:
影响因子:
3.7
通讯作者:
Saito, Priscila T. M.
Saito, Priscila T. M.
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Camargo, Guilherme;Bugatti, Pedro H.;Saito, Priscila T. M.

文献摘要

被引文献

相似文献

由于数据集已经持续增长,已经进行了努力以尝试解决与大量未标记数据与标记数据的稀缺性不成比例相关的问题。另一个重要的问题是,在获得专家提供的注释的困难和需要大量的注释数据,以获得一个强大的分类之间的权衡。在这种情况下,主动学习技术与半监督学习相结合是有趣的。先前选择(通过主动学习策略)并由专家标记的更少量的信息量更大的样本可以将标签传播到一组未标记的数据(通过半监督数据)。然而,大多数的文献作品忽略了交互式响应时间的需要,可以要求某些真实的应用程序。我们提出了一个更有效和高效的主动半监督学习框架,包括一个新的主动学习方法。在生物背景下进行了广泛的实验评估(使用ALL-AML,大肠杆菌和PlantLeaves II数据集),将我们的建议与最先进的文献作品以及不同的监督(SVM,RF,OPF)和半监督(YATSI-SVM,YATSI-RF和YATSI-OPF)分类器进行比较。从所获得的结果中,我们可以观察到我们的框架的好处,它允许分类器更快地实现更高的准确率,同时减少注释样本的数量。此外,我们的主动学习方法所采用的选择标准,基于多样性和不确定性,使学习过程中信息量最大的边界样本的优先级。与其他学习技术相比,我们获得了高达20%的收益。主动半监督学习方法提出了一个更好的权衡(准确性和竞争力和可行的计算时间)相比,主动监督学习的。
Due to datasets have continuously grown, efforts have been performed in the attempt to solve the problem related to the large amount of unlabeled data in disproportion to the scarcity of labeled data. Another important issue is related to the trade-off between the difficulty in obtaining annotations provided by a specialist and the need for a significant amount of annotated data to obtain a robust classifier. In this context, active learning techniques jointly with semi-supervised learning are interesting. A smaller number of more informative samples previously selected (by the active learning strategy) and labeled by a specialist can propagate the labels to a set of unlabeled data (through the semi-supervised one). However, most of the literature works neglect the need for interactive response times that can be required by certain real applications. We propose a more effective and efficient active semi-supervised learning framework, including a new active learning method. An extensive experimental evaluation was performed in the biological context (using the ALL-AML, Escherichia coli and PlantLeaves II datasets), comparing our proposals with state-of-the-art literature works and different supervised (SVM, RF, OPF) and semi-supervised (YATSI-SVM, YATSI-RF and YATSI-OPF) classifiers. From the obtained results, we can observe the benefits of our framework, which allows the classifier to achieve higher accuracies more quickly with a reduced number of annotated samples. Moreover, the selection criterion adopted by our active learning method, based on diversity and uncertainty, enables the prioritization of the most informative boundary samples for the learning process. We obtained a gain of up to 20% against other learning techniques. The active semi-supervised learning approaches presented a better trade-off (accuracies and competitive and viable computational times) when compared with the active supervised learning ones.