Towards Stability of Autoregressive Neural Operators

Towards Stability of Autoregressive Neural Operators
复制标题

DOI:
10.48550/arxiv.2306.10619
复制
发表时间:
2023-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Michael McCabe;P. Harrington;Shashank Subramanian;Jed Brown
Michael McCabe;P. Harrington;Shashank Subramanian;Jed Brown
中科院分区:
其他
文献类型:
--
作者:
Michael McCabe;P. Harrington;Shashank Subramanian;Jed Brown

文献摘要

被引文献

相似文献

神经算子已被证明是物理科学中建模时空系统的一种很有前途的方法。然而,为大型系统训练这些模型可能非常具有挑战性,因为它们会产生大量的计算和内存开销--这些系统通常被迫依赖神经网络的自回归时间步进来预测未来的时间状态。虽然这在管理成本方面是有效的,但随着时间的推移,它可能会导致不受控制的错误增长和最终的不稳定。我们使用物理系统的典型神经算子模型分析了这种自回归误差增长的来源,并探索了缓解这种增长的方法。我们引入了架构和特定于应用程序的改进,允许仔细控制这些模型中导致不稳定的操作,而不会增加计算/内存成本。我们介绍了几个科学系统的结果,包括纳威-斯托克斯流体流动,旋转浅水,和一个高分辨率的全球天气预报系统。我们证明,与这些系统的原始模型相比,将我们的设计原则应用于神经算子可以显著降低长期预测的误差,以及更长的时间范围,而没有定性的发散迹象。为了可重现性,我们将我们的\href{https://github.com/mikemccabe210/stabilizing_neural_operators}{code}开源。
Neural operators have proven to be a promising approach for modeling spatiotemporal systems in the physical sciences. However, training these models for large systems can be quite challenging as they incur significant computational and memory expense -- these systems are often forced to rely on autoregressive time-stepping of the neural network to predict future temporal states. While this is effective in managing costs, it can lead to uncontrolled error growth over time and eventual instability. We analyze the sources of this autoregressive error growth using prototypical neural operator models for physical systems and explore ways to mitigate it. We introduce architectural and application-specific improvements that allow for careful control of instability-inducing operations within these models without inflating the compute/memory expense. We present results on several scientific systems that include Navier-Stokes fluid flow, rotating shallow water, and a high-resolution global weather forecasting system. We demonstrate that applying our design principles to neural operators leads to significantly lower errors for long-term forecasts as well as longer time horizons without qualitative signs of divergence compared to the original models for these systems. We open-source our \href{https://github.com/mikemccabe210/stabilizing_neural_operators}{code} for reproducibility.