Deep forest.

Deep forest.
复制标题

DOI:
10.1093/nsr/nwy108
复制
发表时间:
2019-01
影响因子:
20.6
通讯作者:
Feng J
Feng J
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Zhou ZH;Feng J

文献摘要

参考文献

被引文献

相似文献

目前的深度学习模型大多建立在神经网络的基础上,即可以通过反向传播训练的多层参数可微的非线性模型。在本文中,我们探索了基于不可微模块(如决策树)建立深层模型的可能性。在讨论了深层神经网络背后的奥秘,特别是通过将它们与浅层神经网络和传统的机器学习技术如决策树和Boosting机器进行比较后,我们推测深层神经网络的成功在很大程度上归功于三个特征,即逐层处理、模型内特征转换和足够的模型复杂性。一方面,我们的猜想可能为深度学习的理论理解提供启发;另一方面,为了验证猜想,我们提出了一种产生具有这些特征的深森林的方法。这是一种决策树集成方法,与深度神经网络相比,具有更少的超参数,并且其模型复杂性可以以依赖于数据的方式自动确定。实验表明,该算法对超参数设置具有较强的鲁棒性,在大多数情况下,即使是跨来自不同领域的不同数据,使用相同的默认设置也能获得优异的性能。这项研究打开了基于不可微模块的深度学习的大门,而不需要基于梯度的调整,并展示了构建深度模型而不需要反向传播的可能性。
Current deep-learning models are mostly built upon neural networks, i.e. multiple layers of parameterized differentiable non-linear modules that can be trained by backpropagation. In this paper, we explore the possibility of building deep models based on non-differentiable modules such as decision trees. After a discussion about the mystery behind deep neural networks, particularly by contrasting them with shallow neural networks and traditional machine-learning techniques such as decision trees and boosting machines, we conjecture that the success of deep neural networks owes much to three characteristics, i.e. layer-by-layer processing, in-model feature transformation and sufficient model complexity. On one hand, our conjecture may offer inspiration for theoretical understanding of deep learning; on the other hand, to verify the conjecture, we propose an approach that generates deep forest holding these characteristics. This is a decision-tree ensemble approach, with fewer hyper-parameters than deep neural networks, and its model complexity can be automatically determined in a data-dependent way. Experiments show that its performance is quite robust to hyper-parameter settings, such that in most cases, even across different data from different domains, it is able to achieve excellent performance by using the same default setting. This study opens the door to deep learning based on non-differentiable modules without gradient-based adjustment, and exhibits the possibility of constructing deep models without backpropagation.
DOI: 10.1006/jcss.1997.1504
发表时间: 1997-08-01
影响因子: 1.1
作者:
Freund, Y;Schapire, RE
通讯作者: Schapire, RE
DOI: 10.1007/bf00058655
发表时间: 1996-08-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Breiman, L
通讯作者: Breiman, L
DOI: 10.1023/a:1007682208299
发表时间: 2000-09-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Breiman, L
通讯作者: Breiman, L