Toward a Multi‐Representational Approach to Prediction and Understanding, in Support of Discovery in Hydrology

Toward a Multi‐Representational Approach to Prediction and Understanding, in Support of Discovery in Hydrology
复制标题

DOI:
10.1029/2021wr031548
复制
发表时间:
2022-12
影响因子:
5.4
通讯作者:
Luis De la Fuente;H. Gupta;L. Condon
Luis De la Fuente;H. Gupta;L. Condon
中科院分区:
地球科学1区
文献类型:
--
作者:
Luis De la Fuente;H. Gupta;L. Condon

文献摘要

相似文献

模型开发的关键是选择适当的表示系统,包括对观察到的内容(数据)的表示,以及用于构建输入-状态-输出映射的正式数学结构。这些选择是至关重要的,因为它们完全决定了我们可以提出的问题,我们可以进行的分析和推理的性质,以及我们可以获得的答案。因此,适合于一种调查的代表可能在支持另一种调查的能力方面受到限制。可以说,不同的表征方法如何影响我们从数据中学习的内容还知之甚少。本文探讨了三种代表性的战略,了解流域尺度水文过程如何在水文地质气候多样的智利变化的车辆。具体来说,我们测试了集总水平衡模型(GR 4J)、基于数据的动态系统模型(LSTM)和基于数据的回归树模型(随机森林)。我们获得了关于数据中编码的系统内存、使用替代属性的空间可转移性以及数据集的信息缺陷的见解,这些信息缺陷限制了我们学习适当的输入输出关系的能力。正如预期的那样,每种方法都具有特定的优势,LSTM提供了最好的动态特性,GR 4J在信息不足的条件下最强大,而随机森林回归树方法最支持解释。总体而言,这三种方法的对比性质表明,采用多代表性框架可以更充分地从数据中提取信息,并通过这样做找到更好地实现稳健预测和改善理解目标的信息,最终支持增强的科学发现。
Key to model development is the selection of an appropriate representational system, including both the representation of what is observed (the data), and the formal mathematical structure used to construct the input‐state‐output mapping. These choices are critical, because they completely determine the questions we can ask, the nature of the analyses and inferences we can perform, and the answers we can obtain. Accordingly, a representation that is suitable for one kind of investigation might be limited in its ability to support some other kind. Arguably, how different representational approaches affect what we can learn from data is poorly understood. This paper explores three representational strategies as vehicles for understanding how catchment scale hydrological processes vary across hydro‐geo‐climatologically diverse Chile. Specifically, we test a lumped water‐balance model (GR4J), a data‐based dynamical systems model (LSTM), and a data‐based regression tree model (Random Forest). Insights were obtained regarding system memory encoded in data, spatial transferability by use of surrogate attributes, and informational deficiencies of the data set that limit our ability to learn an adequate input‐output relationship. As expected, each approach exhibits specific strengths, with LSTM providing the best characterization of dynamics, GR4J being the most robust under informationally deficient conditions, and Random Forest regression‐tree method being most supportive of interpretation. Overall, the contrasting nature of the three approaches suggests the value of adopting a multi‐representational framework to more fully extract information from the data and, by doing so, find information that better facilities the goals of robust prediction and improved understanding, ultimately supporting enhanced scientific discovery.