Transferring Hydrologic Data Across Continents – Leveraging Data‐Rich Regions to Improve Hydrologic Prediction in Data‐Sparse Regions

Transferring Hydrologic Data Across Continents – Leveraging Data‐Rich Regions to Improve Hydrologic Prediction in Data‐Sparse Regions
复制标题

DOI:
10.1029/2020wr028600
复制
发表时间:
2021-03
影响因子:
5.4
通讯作者:
K. Ma;D. Feng;K. Lawson;W. Tsai;Chuan Liang;Xiao-rong Huang;Ashutosh Sharma;Chaopeng Shen
K. Ma;D. Feng;K. Lawson;W. Tsai;Chuan Liang;Xiao-rong Huang;Ashutosh Sharma;Chaopeng Shen
中科院分区:
地球科学1区
文献类型:
--
作者:
K. Ma;D. Feng;K. Lawson;W. Tsai;Chuan Liang;Xiao-rong Huang;Ashutosh Sharma;Chaopeng Shen

文献摘要

被引文献

相似文献

现有的全球流量计和集水区属性数据在地理上存在严重的不平衡,数据特征也有很大差异。因此,在一个区域校准的模型通常不能在不作重大修改的情况下迁移到另一个区域。目前,在这些地区,不可转移的机器学习模型习惯于在小型本地数据集上进行训练。在这里,我们展示了迁移学习(TL),在权重初始化和权重冻结的意义上,允许在美国(CONUS,源数据集)上预先训练的长短期记忆(LSTM)径流模型被迁移到其他大陆(目标区域)的集水区,而不需要在目标位置提供广泛的集水区属性。我们证明了这种可能性,数据密集的地区(664个盆地在英国),中等密度(49个盆地在智利中部),和稀缺的遥感属性(5个盆地在中国)。在中国和智利,与使用所有流域的本地训练模型相比,TL模型表现出显着提高的性能。TL的好处随着源数据集中可用数据的数量而增加,并且随着地文多样性的增加而更加明显。TL的好处大于使用未经校准的水文模型输出的预训练LSTM。这些结果表明,世界各地的水文数据具有共性,可以通过深度学习加以利用,并且可以通过简单修改当前工作流程来实现协同效应,从而大大扩展现有大数据的范围。最后,这项工作多样化现有的全球径流基准。
There is a drastic geographic imbalance in available global streamflow gauge and catchment property data, with additional large variations in data characteristics. As a result, models calibrated in one region cannot normally be migrated to another without significant modifications. Currently in these regions, non‐transferable machine learning models are habitually trained over small local data sets. Here we show that transfer learning (TL), in the senses of weight initialization and weight freezing, allows long short‐term memory (LSTM) streamflow models that were pretrained over the conterminous United States (CONUS, the source data set) to be transferred to catchments on other continents (the target regions), without the need for extensive catchment attributes available at the target location. We demonstrate this possibility for regions where data are dense (664 basins in Great Britain), moderately dense (49 basins in central Chile), and scarce with only remotely sensed attributes available (5 basins in China). In both China and Chile, the TL models showed significantly elevated performance compared to locally trained models using all basins. The benefits of TL increased with the amount of available data in the source data set, and seemed to be more pronounced with greater physiographic diversity. The benefits from TL were greater than from pretraining LSTM using the outputs from an uncalibrated hydrologic model. These results suggest hydrologic data around the world have commonalities which could be leveraged by deep learning, and synergies can be had with a simple modification of the current workflows, greatly expanding the reach of existing big data. Finally, this work diversified existing global streamflow benchmarks.