Learning One Convolutional Layer with Overlapping Patches

Learning One Convolutional Layer with Overlapping Patches
复制标题

DOI:
--
复制
发表时间:
2018-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Surbhi Goel;Adam R. Klivans;Raghu Meka
Surbhi Goel;Adam R. Klivans;Raghu Meka
中科院分区:
其他
文献类型:
--
作者:
Surbhi Goel;Adam R. Klivans;Raghu Meka

文献摘要

相似文献

我们给出了第一个可证明有效的算法,用于学习一个隐藏的层卷积网络,相对于一般类(可能重叠的)补丁。此外,我们的算法仅需要对基础分布的温和条件。我们证明,我们的框架捕获了来自计算机视觉的常用方案,包括一维和二维“贴片和大步”卷积。我们的算法($ vertron $)的灵感来自于将等渗回归应用于学习神经网络的最新工作。 VERSOTRON使用一个简单的迭代更新规则,本质上是随机的,对噪声的耐受性(仅要求条件均值函数是一个一层卷积网络,而不是可实现的设置)。与梯度下降相反,韦尔托克斯不需要特殊的初始化或学习速率调整即可融合到全球最佳最佳。我们还指出,学习一个有关高斯分布的隐藏卷积层,只有$一个$ discoint补丁$ p $(其他补丁可能是任意的),在下面的意义上是$简单的$:vervotron可以有效地恢复隐藏的重量通过在$ p $的方向上更新$ $来向量。
We give the first provably efficient algorithm for learning a one hidden layer convolutional network with respect to a general class of (potentially overlapping) patches. Additionally, our algorithm requires only mild conditions on the underlying distribution. We prove that our framework captures commonly used schemes from computer vision, including one-dimensional and two-dimensional "patch and stride" convolutions. Our algorithm-- $Convotron$ -- is inspired by recent work applying isotonic regression to learning neural networks. Convotron uses a simple, iterative update rule that is stochastic in nature and tolerant to noise (requires only that the conditional mean function is a one layer convolutional network, as opposed to the realizable setting). In contrast to gradient descent, Convotron requires no special initialization or learning-rate tuning to converge to the global optimum. We also point out that learning one hidden convolutional layer with respect to a Gaussian distribution and just $one$ disjoint patch $P$ (the other patches may be arbitrary) is $easy$ in the following sense: Convotron can efficiently recover the hidden weight vector by updating $only$ in the direction of $P$.