Differentiable, Learnable, Regionalized Process‐Based Models With Multiphysical Outputs can Approach State‐Of‐The‐Art Hydrologic Prediction Accuracy

Differentiable, Learnable, Regionalized Process‐Based Models With Multiphysical Outputs can Approach State‐Of‐The‐Art Hydrologic Prediction Accuracy
复制标题

DOI:
10.1029/2022wr032404
复制
发表时间:
2022-03
影响因子:
5.4
通讯作者:
D. Feng;Jiangtao Liu;K. Lawson;Chaopeng Shen
D. Feng;Jiangtao Liu;K. Lawson;Chaopeng Shen
中科院分区:
地球科学1区
文献类型:
--
作者:
D. Feng;Jiangtao Liu;K. Lawson;Chaopeng Shen

文献摘要

被引文献

相似文献

整个水循环中水文变量的预测对于水资源管理以及下游应用(如生态系统和水质建模)具有重要价值。最近,像长短期记忆(LSTM)这样的纯数据驱动的深度学习模型在模拟降雨径流和其他地球科学变量方面表现出了看似不可逾越的性能,但它们无法预测未经训练的物理变量,并且仍然难以解释。在这里,我们证明了可微分的、可学习的、基于过程的模型(这里称为δ模型)可以接近LSTM的性能水平,用于区域化参数化的集中观测变量(流量)。我们使用一个简单的水文模型HBV作为主干,并使用嵌入式神经网络(只能在可微编程框架中训练)来参数化、增强或替换基于过程的模型的模块。在不使用集合或后处理器的情况下,δ模型可以获得美国671个流域Daymet强迫数据集的Nash-Sutcliffe效率中位数为0.732,而具有相同设置的最先进LSTM模型为0.748。对于另一个强迫数据集,差异甚至更小:0.715对0.722。同时,由此产生的可学习的基于过程的模型可以输出一整套未经训练的变量,例如土壤和地下水储存量、积雪、蒸散量和基流,并且稍后可以受到其观测结果的约束。模拟的蒸散量和基流排放量与其他估计值一致。通用框架可以处理具有各种过程复杂性的模型,并为从大数据中学习物理开辟了道路。
Predictions of hydrologic variables across the entire water cycle have significant value for water resources management as well as downstream applications such as ecosystem and water quality modeling. Recently, purely data‐driven deep learning models like long short‐term memory (LSTM) showed seemingly insurmountable performance in modeling rainfall runoff and other geoscientific variables, yet they cannot predict untrained physical variables and remain challenging to interpret. Here, we show that differentiable, learnable, process‐based models (called δ models here) can approach the performance level of LSTM for the intensively observed variable (streamflow) with regionalized parameterization. We use a simple hydrologic model HBV as the backbone and use embedded neural networks, which can only be trained in a differentiable programming framework, to parameterize, enhance, or replace the process‐based model's modules. Without using an ensemble or post‐processor, δ models can obtain a median Nash‐Sutcliffe efficiency of 0.732 for 671 basins across the USA for the Daymet forcing data set, compared to 0.748 from a state‐of‐the‐art LSTM model with the same setup. For another forcing data set, the difference is even smaller: 0.715 versus 0.722. Meanwhile, the resulting learnable process‐based models can output a full set of untrained variables, for example, soil and groundwater storage, snowpack, evapotranspiration, and baseflow, and can later be constrained by their observations. Both simulated evapotranspiration and fraction of discharge from baseflow agreed decently with alternative estimates. The general framework can work with models with various process complexity and opens up the path for learning physics from big data.