Online transfer learning with multiple decision trees

Online transfer learning with multiple decision trees
复制标题

具有多个决策树的在线迁移学习

DOI:
10.1007/s13042-019-00998-3
复制
发表时间:
2019
影响因子:
5.6
通讯作者:
Liu Pingshan
Liu Pingshan
中科院分区:
计算机科学3区
文献类型:
--
作者:
Wen Yimin;Qin Yixiu;Qin Keke;Lu Xiaoxia;Liu Pingshan

文献摘要

相似文献

在线学习技术已广泛应用于诸多领域,实例一一呈现。然而,在数据流的早期阶段,在线学习模型无法表现出良好的分类精度,因为它无法收集足够的实例来学习。例如,著名的在线学习算法VFDT需要等待Hoeffding界满足才能进行分裂,这导致数据流开始时的分类精度较差。因此,VFDT可能不适合一些需要快速、准确的在线检测的实际应用。这种情况在存在概念漂移的数据流分类场景中会变得更加严重。本文尝试采用迁移学习算法来弥补VFDT的这一缺点。为了实现这一目标,首先提出了一种名为VFDT-D的新决策树方法,在其叶节点中缓存实例以处理数值属性并适应在线迁移学习(OTL)框架,然后提出考虑树路径、分类精度和分类置信度的度量来评估源域分类器和目标域分类器之间的局部相似性。最后,提出了一种名为 DMOTL 的多源在线迁移学习算法,以 VFDT-D 作为基分类器,并使用所提出的局部相似性度量来选择最佳源域分类器来帮助迁移学习。对几个合成和真实数据集的广泛实验证明了所提出算法的优势。
Online learning techniques have been widely used in many fields where instances come one by one. However, in early stage of a data stream, online learning models cannot exhibit good classification accuracy for it cannot collect sufficient instances to learn. For example, a well-known online learning algorithm named as very fast decision tree (VFDT) needs to wait for Hoeffding bound satisfied to split, which leads to poor classification accuracy at the beginning of data stream. Thus, VFDT may not be appropriate for some real applications which demand us a fast and accurate online detection. This situation will become more serious in the scenario of data stream classification with concept drift. This paper attempts to take transfer learning algorithm to make up this shortcoming of VFDT. To achieve this goal, a new decision tree method named as VFDT-D is first proposed to cache instances in its leaf nodes to handle numerical attributes and adapt to a framework of online transfer learning (OTL), and then a measure which considers tree path, classification accuracy and classification confidence is proposed to evaluate the local similarity between source and target domain classifiers. At last, a multiple-source online transfer learning algorithm named as DMOTL is proposed to take VFDT-D as base classifier and use the proposed measure of local similarity to select the optimal source domain classifier to help transfer learning. The extensive experiments on several synthetic and real-world datasets demonstrate the advantage of the proposed algorithm.