Extreme Learning Machine for Multilayer Perceptron

Extreme Learning Machine for Multilayer Perceptron
复制标题

多层感知器的极限学习机

DOI:
10.1109/tnnls.2015.2424995
复制
发表时间:
2016-04-01
影响因子:
10.4
通讯作者:
Huang, Guang-Bin
Huang, Guang-Bin
中科院分区:
计算机科学1区
文献类型:
--
作者:
Tang, Jiexiong;Deng, Chenwei;Huang, Guang-Bin

文献摘要

被引文献

相似文献

极限学习机(ELM)是一种新兴的广义单隐层前馈神经网络学习算法,其隐含节点参数是随机生成的,输出权重是解析计算的。然而,由于其浅层架构,使用 ELM 进行特征学习可能对自然信号(例如图像/视频)无效,即使有大量隐藏节点也是如此。为了解决这个问题,本文提出了一种新的基于 ELM 的多层感知器分层学习框架。所提出的架构分为两个主要部分:1)自学特征提取,然后是监督特征分类;2)它们通过随机初始化的隐藏权重进行桥接。本文的新颖之处如下:1)采用无监督多层编码进行特征提取,并通过ℓ1约束开发基于ELM的稀疏自动编码器。通过这样做,它实现了比原始 ELM 更紧凑、更有意义的特征表示; 2)利用ELM随机特征映射的优势,在最终决策之前对分层编码的输出进行随机投影,从而获得更好的泛化能力和更快的学习速度; 3)与深度学习(DL)的贪婪分层训练不同,所提出框架的隐藏层以前向方式进行训练。一旦前一层建立起来,当前层的权重就固定不变,无需微调。因此,它的学习效率比DL要好得多。对各种广泛使用的分类数据集的大量实验表明,所提出的算法比现有最先进的分层学习方法实现了更好、更快的收敛。此外,计算机视觉中的多种应用进一步证实了所提出的学习方案的通用性和能力。
Extreme learning machine (ELM) is an emerging learning algorithm for the generalized single hidden layer feedforward neural networks, of which the hidden node parameters are randomly generated and the output weights are analytically computed. However, due to its shallow architecture, feature learning using ELM may not be effective for natural signals (e.g., images/videos), even with a large number of hidden nodes. To address this issue, in this paper, a new ELM-based hierarchical learning framework is proposed for multilayer perceptron. The proposed architecture is divided into two main components: 1) self-taught feature extraction followed by supervised feature classification and 2) they are bridged by random initialized hidden weights. The novelties of this paper are as follows: 1) unsupervised multilayer encoding is conducted for feature extraction, and an ELM-based sparse autoencoder is developed via ℓ1 constraint. By doing so, it achieves more compact and meaningful feature representations than the original ELM; 2) by exploiting the advantages of ELM random feature mapping, the hierarchically encoded outputs are randomly projected before final decision making, which leads to a better generalization with faster learning speed; and 3) unlike the greedy layerwise training of deep learning (DL), the hidden layers of the proposed framework are trained in a forward manner. Once the previous layer is established, the weights of the current layer are fixed without fine-tuning. Therefore, it has much better learning efficiency than the DL. Extensive experiments on various widely used classification data sets show that the proposed algorithm achieves better and faster convergence than the existing state-of-the-art hierarchical learning methods. Furthermore, multiple applications in computer vision further confirm the generality and capability of the proposed learning scheme.