课题基金 / 基金详情

Combining Machine Learning and Data Assimilation to infer model errors

Combining Machine Learning and Data Assimilation to infer model errors
结合机器学习和数据同化来推断模型错误
批准号:
2267924
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
数据同化(DA)是将观测(或“数据”)中包含的信息与现有系统的先验知识相结合的科学,通常以一组耦合的偏微分方程的形式出现。它是贝叶斯推理应用于地球科学,特别是气象学和气候科学,在那里流体动力学的规律是已知的和必须使用。然后将这些知识转化为随时间变化的数值模型,模型的大小远远大于可用数据的数量。由于这种模型与观测的维度不匹配,模型对数据分析的结果起着关键作用,因此可以将其视为模型驱动的过程。数据分析中使用的地流体数值模式只是真实大气、海洋或整个气候系统的近似表示。必须考虑到由此产生的模型误差,并且已经投入了大量的工作来使数据分析方法能够以统计方式容纳模型误差。真正的模型误差是未知的,它由许多不同的来源引起,如数值离散化、参数误差和未解析尺度的存在。特别是子网格过程对模型的技能非常关键,并在适当的子网格参数化方案中进行描述。估计这些格式的形式及其参数的值对于预测和成功的数据同化都是至关重要的。现有的估计技术在很大程度上依赖于物理直觉和对现有观测结果的特别使用。然而,目前还没有系统、可靠的方法来估计模型与观测数据失配的模型误差。另一方面,近年来可观测数据的不断增加,加上计算能力的惊人增长,使得完全数据驱动的方法成为可能。这种数据驱动的革命主要是由机器学习(ML)技术(例如深度神经网络等)的蓬勃发展推动的,这些技术越来越成功地显示出能够从多元数据集中提取潜在的动态规律,具有令人印象深刻的预测技能和对复杂行为进行分类的能力。目前,ML和DA算法非常相似:两种方法都在给定一组目标(即观察值)的情况下优化参数。优化,或ML术语中的训练,需要在DA中计算梯度和伴随,在ML中称为反向传播。主要区别在于,在DA中,模型被明确地设置为一组物理c约束,而在ML中,直到最近才成熟,以纳入我们的物理知识。此外,与数据分析相反,机器学习和数据分析的互补性,数据分析在地球科学中的成功,以及机器学习在同一领域的美好未来,促使人们寻找合适的组合,充分利用它们的优势,减轻它们的弱点。拟议的博士研究项目将在这个边界上工作。我们建议使用机器学习来“学习”核心动态模型中未明确描述的子网格过程的参数化,基于来自DA实验的模型-观测不匹配数据。然后,这个新的参数化将在模型中实现并用于执行数据分析。由于数据分析性能与所使用的模型误差呈非线性关系,由此产生的模型观测不匹配数据可以再次被ML用于改进其参数化描述,然后可以在数据分析中使用。这个迭代过程,如果定义良好,将收敛,通过模型改进和模型的高级DA/ML初始化,导致环境预测的潜在突破。
英文摘要
Data assimilation (DA) is the science of combining information contained in observations (or 'data') with prior knowledge of the system at hand, typically in the form of a coupled set of partial differential equations. It is Bayesian Inference applied to the geosciences, especially meteorology and climate science, where the laws of fluid-dynamics are known and bound to be used. This knowledge is then translated into a time-evolving numerical model and the size of the model is vastly larger than the amount of available data. As a consequence of this model-to-observation dimensional mismatch, the models play a critical role on the outcome of DA, that can thus be seen as a model-driven procedure. Numerical models of geo-fluids used in DA are only an approximate representation of the real atmosphere, ocean or the whole climate system. The resulting model error has to be taken into account and a substantial amount of work has been devoted to make DA methods able to accommodate model error in a statistical way. The real model error is unknown, and arises from many different sources, such as numerical discretization, parametric error, and the presence of unresolved scales. Sub-grid processes in particular are very critical to the skill of the models and are described in appropriate sub-grid parametrization schemes. Estimating the form of these schemes and the values of their parameters is of crucial importance, both for prediction and for successful data assimilation. Existing estimation techniques rely heavily on physical intuition and ad-hoc use of available observations. However, systematic and robust methods to estimate model errors from model-observation mismatch data do not exist.On the other hand, in recent times the constant increase of available observations, accompanied by a similarly spectacular growth of the computing power, have made fully data-driven approaches possible. This data-driven revolution has been mostly pushed by the flourishing of machine-learning (ML) techniques (e.g. deep neural networks, among others) that with increasing success have shown to be able to extract the underlying dynamical laws from a multivariate dataset, with impressive predictive skill and capabilities to classify complex behaviors.At present, ML and DA algorithms are quite similar: both approaches optimize parameters given a set of targets (i.e., the observations). The optimization, or the training in ML jargon, requires computing gradients and adjoints in DA, referred to as backpropagation in ML. The major difference is that while in the DA the model is explicitly set out as a set of physical c constraints, in ML is only recently maturing to incorporate our physical knowledge. Furthermore, as opposed to DA, no principled uncertainty quantification is used in ML.The complementarities of ML and DA, the success of DA in the geoscience, and the promising future of ML in the same area, motivates the search for suitable combinations of them that adequately exploit each of their strength and mitigate each of their weaknesses. The proposed PhD research program is will work at this boundary.We propose to use machine learning to "learn" the parametrization of sub-grid processes that are not explicitly described in the core dynamical model, based on model-observation mismatch data from DA experiments. This new parametrization will then be implemented in the model and used to perform DA. Since the DA performance is nonlinearly related to the model error used, the resulting model-observation mismatch data can again we used by ML to improve its parameterization description, which can then be used in DA. This iterative process, if well defined, will converge, leading to a potential breakthrough in environmental prediction via both model improvement and superior DA/ML initialization of the models.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位: