Non-asymptotic and Accurate Learning of Nonlinear Dynamical Systems

Non-asymptotic and Accurate Learning of Nonlinear Dynamical Systems
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Yahya Sattar;Samet Oymak
Yahya Sattar;Samet Oymak
中科院分区:
其他
文献类型:
--
作者:
Yahya Sattar;Samet Oymak

文献摘要

被引文献

相似文献

研究了非线性状态方程h_{t +1}=\phi(h_t,u_t;\theta)+w_t $所控制的可镇定系统的学习问题。这里$\theta $是未知的系统动态,$h_t $是状态,$u_t $是输入,$w_t $是加性噪声向量。我们研究基于梯度的算法来学习系统动力学$\theta $从一个单一的有限轨迹获得的样本。如果系统是由一个稳定的输入政策运行,我们表明,时间依赖的样本可以近似由独立同分布。使用混合时间参数通过截断参数进行采样。然后,我们开发新的保证经验损失的梯度的一致收敛。与现有的工作不同,我们的边界是噪声敏感的,这允许以高精度和小样本复杂度学习地面实况动态。总之,我们的研究结果有利于有效的学习一般非线性系统下稳定的政策。我们专门保证入境明智的非线性激活和验证我们的理论在各种数值实验
We consider the problem of learning stabilizable systems governed by nonlinear state equation $h_{t+1}=\phi(h_t,u_t;\theta)+w_t$. Here $\theta$ is the unknown system dynamics, $h_t $ is the state, $u_t$ is the input and $w_t$ is the additive noise vector. We study gradient based algorithms to learn the system dynamics $\theta$ from samples obtained from a single finite trajectory. If the system is run by a stabilizing input policy, we show that temporally-dependent samples can be approximated by i.i.d. samples via a truncation argument by using mixing-time arguments. We then develop new guarantees for the uniform convergence of the gradients of empirical loss. Unlike existing work, our bounds are noise sensitive which allows for learning ground-truth dynamics with high accuracy and small sample complexity. Together, our results facilitate efficient learning of the general nonlinear system under stabilizing policy. We specialize our guarantees to entry-wise nonlinear activations and verify our theory in various numerical experiments