Effects of Nonlinearity and Network Architecture on the Performance of Supervised Neural Networks

Effects of Nonlinearity and Network Architecture on the Performance of Supervised Neural Networks
复制标题

DOI:
10.3390/a14020051
复制
发表时间:
2021-02-01
期刊:
影响因子:
2.3
通讯作者:
Wang, Yunjiao
Wang, Yunjiao
中科院分区:
其他
文献类型:
--
作者:
Kulathunga, Nalinda;Ranasinghe, Nishath Rajiv;Wang, Yunjiao

文献摘要

被引文献

相似文献

深度学习模型中使用的激活函数的非线性对于预测模型的成功至关重要。几个简单的非线性函数,包括整流线性单元(ReLU)和泄漏-ReLU(L-ReLU),通常用于神经网络中以施加非线性。在实际应用中,这些功能显著提高了模型的精度。然而,有有限的洞察神经网络中的非线性对其性能的影响。在这里,我们在不同的模型架构和数据域的背景下,使用ReLU和L-ReLU激活函数来研究神经网络模型作为非线性函数的性能。我们使用熵作为随机性的度量,来量化不同结构形状的非线性对神经网络性能的影响。我们表明,当网络具有足够数量的参数时,ReLU非线性是激活函数的更好选择。然而,我们发现具有迁移学习的图像分类模型似乎在全连接层中使用L-ReLU表现良好。我们发现,在神经网络的隐层输出的熵可以公平地代表信息损失的非线性函数的波动。此外,我们调查的熵轮廓的浅层神经网络的一种方式来表示其隐藏层的动态。
The nonlinearity of activation functions used in deep learning models is crucial for the success of predictive models. Several simple nonlinear functions, including Rectified Linear Unit (ReLU) and Leaky-ReLU (L-ReLU) are commonly used in neural networks to impose the nonlinearity. In practice, these functions remarkably enhance the model accuracy. However, there is limited insight into the effects of nonlinearity in neural networks on their performance. Here, we investigate the performance of neural network models as a function of nonlinearity using ReLU and L-ReLU activation functions in the context of different model architectures and data domains. We use entropy as a measurement of the randomness, to quantify the effects of nonlinearity in different architecture shapes on the performance of neural networks. We show that the ReLU nonliearity is a better choice for activation function mostly when the network has sufficient number of parameters. However, we found that the image classification models with transfer learning seem to perform well with L-ReLU in fully connected layers. We show that the entropy of hidden layer outputs in neural networks can fairly represent the fluctuations in information loss as a function of nonlinearity. Furthermore, we investigate the entropy profile of shallow neural networks as a way of representing their hidden layer dynamics.