A flexible and efficient knowledge-guided machine learning data assimilation (KGML-DA) framework for agroecosystem prediction in the US Midwest

A flexible and efficient knowledge-guided machine learning data assimilation (KGML-DA) framework for agroecosystem prediction in the US Midwest
复制标题

DOI:
10.1016/j.rse.2023.113880
复制
发表时间:
2023-12
影响因子:
13.5
通讯作者:
Qi Yang;Licheng Liu;Junxiong Zhou;Rahul Ghosh;Bin Peng;Kaiyu Guan;Jinyun Tang;Wang Zhou
Qi Yang;Licheng Liu;Junxiong Zhou;Rahul Ghosh;Bin Peng;Kaiyu Guan;Jinyun Tang;Wang Zhou
中科院分区:
工程技术1区
文献类型:
--
作者:
Qi Yang;Licheng Liu;Junxiong Zhou;Rahul Ghosh;Bin Peng;Kaiyu Guan;Jinyun Tang;Wang Zhou

文献摘要

被引文献

相似文献

基于过程的模型被广泛用于农业生态系统动态预测,但由于模型结构不完善、模型参数有偏差以及模型输入不准确或难以接近,这种模型结果往往含有相当大的不确定性。数据同化(Data assimilation, DA)技术被广泛采用,通过校正模型参数或利用观测值动态更新模型状态变量来降低预测的不确定性。然而,计算成本高、模型结构误差难以缓解、框架开发灵活性低等问题阻碍了其在大规模农业生态系统预测中的应用。在本研究中,我们通过提出一种新的数据分析框架来解决这些挑战,该框架将基于知识引导的机器学习(KGML)的代理与张张集成卡尔曼滤波器(EnKF)和并行粒子群优化(PSO)相结合,以有效地吸收历史和季节性多源遥感数据。具体来说,我们将基于过程的模型ecosys中的知识整合到基于门控循环单元(GRU)的分层神经网络中。KGML-DA的层次结构模拟了生态系统的关键过程,并在目标变量之间建立了因果关系。以美国玉米带的碳预算量化为背景,我们评估了KGML-DA在三个农业样点(US- ne1、US- ne2、US- ne3)、县级(627个县)和30 m像素级(伊利诺伊州尚佩恩县)粮食产量预测碳循环关键过程中的表现。现场试验结果表明,更新初级生产总值(GPP)等上游变量可以改善下游生态系统呼吸、净生态系统交换、生物量和叶面积指数(LAI)等变量的预测,玉米的RMSE降低9.2% ~ 30.5%,大豆的RMSE降低4.8% ~ 24.6%。在修正上游变量后,下游变量的不确定性被自动约束,证明了层次代理中因果关系的有效性。研究发现,季内GPP和蒸散发(ET)产品与历史GPP和调查产量联合使用对县域产量的预测效果最好,而同化季内LAI观测值有利于极端年份的预测。区域产量估算的不确定性和误差分析表明,KGML-DA可将玉米和大豆的预测误差分别降低26.5%和36.2%。值得注意的是,基于gpu的张量运算设计使该数据分析框架比具有高性能计算系统的PB模型快7000倍以上,这表明所提出的框架在季节性、高分辨率农业生态系统预测方面具有很高的潜力。
Process-based models are widely used to predict the agroecosystem dynamics, but such modeled results often contain considerable uncertainty due to the imperfect model structure, biased model parameters, and inaccurate or inaccessible model inputs. Data assimilation (DA) techniques are widely adopted to reduce prediction uncertainty by calibrating model parameters or dynamically updating the model state variables using observations. However, high computational cost, difficulties in mitigating model structural error, and low flexibility in framework development hinder its applications in large-scale agroecosystem predictions. In this study, we addressed these challenges by proposing a novel DA framework that integrates a Knowledge-Guided Machine Learning (KGML)-based surrogate with tensorized ensemble Kalman filter (EnKF) and parallelized particle swarm optimization (PSO) to effectively assimilate historical and in-season multi-source remote sensing data. Specifically, we incorporate knowledge from a process-based model,ecosys, into a Gated Recurrent Unit (GRU)-based hierarchical neural network. The hierarchical architecture of KGML-DA mimics key processes ofecosysand builds a causal relationship between target variables. Using carbon budget quantification in the US Corn-Belt as a context, we evaluated KGML-DA's performance in predicting key processes of the carbon cycle at three agricultural sites (US-Ne1, US-Ne2, US-Ne3), along with county-level (627 counties) and 30-m pixel-level (Champaign County, IL) grain yield. The site experiments show that updating the upstream variable, e.g., gross primary production (GPP), improved the prediction of downstream variables such as ecosystem respiration, net ecosystem exchange, biomass, and leaf area index (LAI), with RMSE reductions ranging from 9.2% to 30.5% for corn and 4.8% to 24.6% for soybean. Uncertainty in downstream variables was automatically constrained after correcting the upstream variables, demonstrating the effectiveness of the causality linkages in the hierarchical surrogate. We found joint use of in-season GPP and evapotranspiration (ET) products along with historical GPP and surveyed yields achieved the best prediction for county-level yields, while assimilating in-season LAI observations benefitted the prediction in extreme years. Uncertainty and error analysis of regional yield estimation demonstrated that KGML-DA could reduce prediction error by 26.5% for corn and 36.2% for soybean. Remarkably, the GPU-based tensor operation design makes this DA framework more than 7000 times faster than the PB model with a High-Performance Computing system, indicating the high potential of the proposed framework for in-season, high-resolution agroecosystem predictions.