Machine Learning: Deepest Learning as Statistical Data Assimilation Problems

Machine Learning: Deepest Learning as Statistical Data Assimilation Problems
复制标题

DOI:
10.1162/neco_a_01094
复制
发表时间:
2018-08-01
期刊:
影响因子:
2.9
通讯作者:
Shirman, Sasha
Shirman, Sasha
中科院分区:
计算机科学4区
文献类型:
--
作者:
Abarbanel, Henry D., I;Rozdeba, Paul J.;Shirman, Sasha

文献摘要

被引文献

相似文献

我们制定了机器学习和制定统计数据同化广泛应用于物理和生物科学之间的等价关系。对应关系是前馈人工网络设置中的层数是数据同化设置中的时间的模拟。在机器学习文献中已经注意到这种联系。我们增加了一个视角,扩展了统计物理学方法以及拉格朗日和哈密顿动力学方面如何在网络训练和设计中发挥作用。在这种等价性的讨论中,我们表明,增加更多的层(使网络更深)类似于在数据同化框架中增加时间分辨率。我们还讨论了将这种等价性扩展到递归网络的问题。我们探讨了如何使用数据同化的方法在机器学习环境中找到全局最小代价函数的候选者。计算简单的模型,从双方的等价reported.Also讨论的是一个框架,其中的时间或层标签是连续的,提供了一个微分方程,欧拉-拉格朗日方程及其边界条件,作为必要条件的最小值的成本函数。这表明,所解决的问题是一个两点边值问题熟悉的讨论变分方法。连续层的使用被称为“最深学习”。“这些问题在连续层相空间中具有辛对称性。拉格朗日版本和哈密顿版本的这些问题。他们的研究实施在一个离散的时间/层,同时尊重辛结构,解决。哈密顿版本提供了一个直接的理由,反向传播作为一种解决方法,为一定的两点边值问题。
We formulate an equivalence between machine learning and the formulation of statistical data assimilation as used widely in physical and biological sciences. The correspondence is that layer number in a feedforward artificial network setting is the analog of time in the data assimilation setting. This connection has been noted in the machine learning literature. We add a perspective that expands on how methods from statistical physics and aspects of Lagrangian and Hamiltonian dynamics play a role in how networks can be trained and designed. Within the discussion of this equivalence, we show that adding more layers (making the network deeper) is analogous to adding temporal resolution in a data assimilation framework. Extending this equivalence to recurrent networks is also discussed.We explore how one can find a candidate for the global minimum of the cost functions in the machine learning context using a method from data assimilation. Calculations on simple models from both sides of the equivalence are reported.Also discussed is a framework in which the time or layer label is taken to be continuous, providing a differential equation, the Euler-Lagrange equation and its boundary conditions, as a necessary condition for a minimum of the cost function. This shows that the problem being solved is a two-point boundary value problem familiar in the discussion of variational methods. The use of continuous layers is denoted "deepest learning."These problems respect a symplectic symmetry in continuous layer phase space. Both Lagrangian versions and Hamiltonian versions of these problems are presented. Their well-studied implementation in a discrete time/layer, while respecting the symplectic structure, is addressed. The Hamiltonian version provides a direct rationale for backpropagation as a solution method for a certain two-point boundary value problem.