On a Fitting of a Heaviside Function by Deep ReLU Neural Networks

On a Fitting of a Heaviside Function by Deep ReLU Neural Networks
复制标题

深度 ReLU 神经网络对 Heaviside 函数的拟合

DOI:
10.1007/978-3-030-04167-0_6
复制
发表时间:
2018
期刊:
L. Cheng et al. (Eds.): ICONIP 2018, LNCS 11301
影响因子:
--
通讯作者:
Hagiwara Katsuyuki
Hagiwara Katsuyuki
中科院分区:
--
文献类型:
--
作者:
Kanako Komiya;Takumi Seitou;Minoru Sasaki;Hiroyuki Shinnou;辰己晶洋,青木なつみ,今西康仁,升元颯人,中川皓登,大久保雅史;Hagiwara Katsuyuki

文献摘要

相似文献

最近对深度神经网络的研究兴趣是理解为什么深度网络比浅网络更受欢迎。在这篇文章中,我们考虑了深层结构在实现训练中的并行功能方面的优势。这不仅是简单的分类问题,而且是构造一般非光滑函数的基础。如果我们可以设置非常大的权重值,则可以通过ReLU的差来很好地近似一个单边函数。然而,要在训练中做到这一点并不容易。我们证明了,如果我们采用深层结构,则可以在没有大的权重值的情况下很好地表示单侧函数。我们还表明,如果一个网络被训练来实现一个单边函数,那么输入端的权值更新项必然很大。因此,通过设置较小的学习率,可以明显加快训练速度。因此,我们可以说,通过采用深层结构,可以在适度小的学习率下,在合理的训练时间内获得良好的拟合。我们的研究结果表明,一个深的结构是有效的,在一个实际的训练,需要一个不连续的输出。
A recent research interest on deep neural networks is to understand why deep networks are preferred to shallow networks. In this article, we considered an advantage of a deep structure in realizing a heaviside function in training. This is significant not only as simple classification problems but also as a basis in constructing general non-smooth functions. A heaviside function can be well approximated by a difference of ReLUs if we can set extremely large weight values. However, it is not so easy to attain them in training. We showed that a heaviside function can be well represented without large weight values if we employ a deep structure. We also showed that update terms of weights at input side can be necessarily large if a network is trained to realize a heaviside function. Therefore, apparent acceleration of training is brought about by setting a small learning rate. As a result, we can say that, by employing a deep structure, a good fitting of heaviside function can be obtained within a reasonable training time under a moderate small learning rate. Our results suggest that a deep structure is effective in a practical training that requires a discontinuous output.