Learning One-hidden-layer Neural Networks under General Input Distributions

Learning One-hidden-layer Neural Networks under General Input Distributions
复制标题

DOI:
--
复制
发表时间:
2018-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Weihao Gao;Ashok Vardhan Makkuva;Sewoong Oh;P. Viswanath
Weihao Gao;Ashok Vardhan Makkuva;Sewoong Oh;P. Viswanath
中科院分区:
其他
文献类型:
--
作者:
Weihao Gao;Ashok Vardhan Makkuva;Sewoong Oh;P. Viswanath

文献摘要

被引文献

相似文献

最近在训练神经网络上取得了重大进展,在训练神经网络上,主要的挑战是解决了具有丰富关键点的优化问题。但是,现有的解决此问题的方法取决于限制性假设:培训数据是从高斯分布中汲取的。在本文中,我们为设计损失功能提供了一种新颖的统一框架,该损失功能具有理想的景观特性,可用于广泛的一般输入分布。在这些损失函数上,显着地,理论上,随机梯度下降以全局初始化恢复了真实参数,并在经验上优于现有方法。我们的损失功能设计桥接了分数功能的概念,即神经网络优化的主题。我们方法的核心是从样本中估算得分函数的任务,这是理论统计的基本和独立的兴趣。传统估计方法(例如:基于内核)一开始就失败;我们带来局部可能性的统计方法来设计一个新的分数函数估计量,该方法可适应未知密度的局部几何形状。
Significant advances have been made recently on training neural networks, where the main challenge is in solving an optimization problem with abundant critical points. However, existing approaches to address this issue crucially rely on a restrictive assumption: the training data is drawn from a Gaussian distribution. In this paper, we provide a novel unified framework to design loss functions with desirable landscape properties for a wide range of general input distributions. On these loss functions, remarkably, stochastic gradient descent theoretically recovers the true parameters with global initializations and empirically outperforms the existing approaches. Our loss function design bridges the notion of score functions with the topic of neural network optimization. Central to our approach is the task of estimating the score function from samples, which is of basic and independent interest to theoretical statistics. Traditional estimation methods (example: kernel based) fail right at the outset; we bring statistical methods of local likelihood to design a novel estimator of score functions, that provably adapts to the local geometry of the unknown density.