Convex and concave envelopes of artificial neural network activation functions for deterministic global optimization

Convex and concave envelopes of artificial neural network activation functions for deterministic global optimization
复制标题

DOI:
10.1007/s10898-022-01228-x
复制
发表时间:
2022-08
影响因子:
1.8
通讯作者:
Matthew E. Wilhelm;Chenyu Wang;M. D. Stuber
Matthew E. Wilhelm;Chenyu Wang;M. D. Stuber
中科院分区:
数学3区
文献类型:
--
作者:
Matthew E. Wilhelm;Chenyu Wang;M. D. Stuber

文献摘要

被引文献

相似文献

在这项工作中,我们提出了一般的方法来构建凸/凹松弛的激活函数,通常选择人工神经网络(ANN)。这些函数的选择通常是由更广泛的建模考虑因素与高计算性能的需求平衡来决定的。直接应用可分解规划技术来计算这些函数的边界和凸/凹松弛往往会导致由于依赖性问题的弱外壳。此外,分段公式,定义了几个流行的激活函数,防止计算的凸/凹松弛,因为他们违反了因式分解功能的要求。为了提高确定性全局优化应用的神经网络松弛的性能,本研究提出了一个图书馆的信封的深入研究整流器型和sigmoid激活函数的发展,除了新的self-gated sigmoid加权线性单元(SiLU)和高斯误差线性单元激活函数。我们证明了激活函数的包络直接导致ANN在其输入域上的更紧松弛。反过来,这些改进转化为CPU运行时间的显着减少所需的解决优化问题,涉及人工神经网络模型的ε-全局最优。我们进一步证明,可分解的编程方法导致上级的计算性能比其他国家的最先进的方法。
In this work, we present general methods to construct convex/concave relaxations of the activation functions that are commonly chosen for artificial neural networks (ANNs). The choice of these functions is often informed by both broader modeling considerations balanced with a need for high computational performance. The direct application of factorable programming techniques to compute bounds and convex/concave relaxations of such functions often lead to weak enclosures due to the dependency problem. Moreover, the piecewise formulation that defines several popular activation functions, prevents the computation of convex/concave relaxations as they violate the factorable function requirement. To improve the performance of relaxations of ANNs for deterministic global optimization applications, this study presents the development of a library of envelopes of the thoroughly studied rectifier-type and sigmoid activation functions, in addition to the novel self-gated sigmoid-weighted linear unit (SiLU) and Gaussian error linear unit activation functions. We demonstrate that the envelopes of activation functions directly lead to tighter relaxations of ANNs on their input domain. In turn, these improvements translate to a dramatic reduction in CPU runtime required for solving optimization problems involving ANN models to epsilon-global optimality. We further demonstrate that the factorable programming approach leads to superior computational performance over alternative state-of-the-art approaches.