Parallel Deep Neural Networks Have Zero Duality Gap

Parallel Deep Neural Networks Have Zero Duality Gap
复制标题

并行深度神经网络具有零对偶间隙

DOI:
--
复制
发表时间:
2021
期刊:
International Conference on Learning Representations
影响因子:
--
通讯作者:
Mert Pilanci
Mert Pilanci
中科院分区:
--
文献类型:
--
作者:
Yifei Wang;Tolga Ergen;Mert Pilanci

文献摘要

参考文献

被引文献

相似文献

培训深度神经网络是一个具有挑战性的非凸优化问题。最近的工作证明,正规化有限宽度的两层relu网络的强双重性(这意味着零双重性差距),因此提供了同等的凸训练问题。但是,将此结果扩展到更深的网络仍然是一个开放的问题。在本文中,我们证明具有向量输出的更深线性网络的二元性差距为非零。相比之下,我们表明可以通过并行堆叠标准深网络来获得零二元性差距,我们称之为并行体系结构并修改正则化。因此,我们证明了等效凸问题的强大双重性和存在,这些问题能够对深层网络进行全球最佳培训。作为我们分析的副产品,我们证明了网络参数的权重衰减正则化明确地通过封闭形式表达式鼓励低级解决方案。此外,我们表明,给定级别1数据矩阵的三层标准relu网络的强双重性具有强度。
Training deep neural networks is a challenging non-convex optimization problem. Recent work has proven that the strong duality holds (which means zero duality gap) for regularized finite-width two-layer ReLU networks and consequently provided an equivalent convex training problem. However, extending this result to deeper networks remains to be an open problem. In this paper, we prove that the duality gap for deeper linear networks with vector outputs is non-zero. In contrast, we show that the zero duality gap can be obtained by stacking standard deep networks in parallel, which we call a parallel architecture, and modifying the regularization. Therefore, we prove the strong duality and existence of equivalent convex problems that enable globally optimal training of deep networks. As a by-product of our analysis, we demonstrate that the weight decay regularization on the network parameters explicitly encourages low-rank solutions via closed-form expressions. In addition, we show that strong duality holds for three-layer standard ReLU networks given rank-1 data matrices.
DOI: 10.48550/arxiv.2205.08078
发表时间: 2022-05
期刊: ArXiv
影响因子: --
作者:
Arda Sahiner;Tolga Ergen;Batu Mehmet Ozturkler;J. Pauly;M. Mardani;Mert Pilanci
通讯作者: Arda Sahiner;Tolga Ergen;Batu Mehmet Ozturkler;J. Pauly;M. Mardani;Mert Pilanci
DOI: --
发表时间: 2020-07
期刊: ArXiv
影响因子: --
作者:
E. Moroshko;Suriya Gunasekar;Blake E. Woodworth;J. Lee;N. Srebro;Daniel Soudry
通讯作者: E. Moroshko;Suriya Gunasekar;Blake E. Woodworth;J. Lee;N. Srebro;Daniel Soudry
超参数化神经网络的凸几何和对偶性
DOI: --
发表时间: 2021
影响因子: 6
作者:
Ergen, T.;Pilanci, M.
通讯作者: Pilanci, M.
揭秘 ReLU 网络中的批量归一化:等效凸优化模型和隐式正则化
DOI: --
发表时间: 2022
期刊: International Conference on Learning Representations
影响因子: --
作者:
Ergen, T.;Sahiner, A.;Ozturkler, B.;Pauly, J.;Mardani, M.;Pilanci, M.
通讯作者: Pilanci, M.
DOI: --
发表时间: 2020-12
期刊: ArXiv
影响因子: --
作者:
Arda Sahiner;Tolga Ergen;J. Pauly;Mert Pilanci
通讯作者: Arda Sahiner;Tolga Ergen;J. Pauly;Mert Pilanci