A simulation study of disaggregation regression for spatial disease mapping

A simulation study of disaggregation regression for spatial disease mapping
复制标题

DOI:
10.1002/sim.9220
复制
发表时间:
2021-10-17
影响因子:
2
通讯作者:
Cameron, Ewan
Cameron, Ewan
中科院分区:
医学3区
文献类型:
--
作者:
Arambepola, Rohan;Lucas, Tim C. D.;Cameron, Ewan

文献摘要

被引文献

相似文献

分解回归已成为空间疾病绘图中的重要工具,用于根据汇总的响应数据对疾病风险进行精细预测。通过包含高分辨率协变量信息并在精细尺度上对数据生成过程进行建模,希望这些模型能够在精细空间尺度上准确地学习协变量和响应之间的关系。然而,验证这些高分辨率预测可能是一个挑战,因为通常在这个空间尺度上没有观察到数据。在本研究中,对各种设置下的模拟数据进行分解回归,并将所得的精细预测与模拟的地面实况进行比较。使用不同数量的数据点、聚合区域的大小和模型错误指定的级别来调查性能。还研究了总体水平交叉验证作为精细预测性能衡量标准的有效性。随着观测数量的增加和聚合区域大小的减小,预测性能得到提高。当模型被明确指定时,即使观测数量较少且聚合区域较大,精细尺度的预测也是准确的。在模型错误指定的情况下,大聚合区域的预测性能明显较差,但当响应数据聚合在较小区域时,预测性能仍然很高。总体水平上的交叉验证相关性是精细尺度预测性能的良好预测指标。虽然这些模拟不太可能捕捉现实生活中响应数据的细微差别,但这项研究深入了解了不同背景下分解回归的有效性。
Disaggregation regression has become an important tool in spatial disease mapping for making fine-scale predictions of disease risk from aggregated response data. By including high resolution covariate information and modeling the data generating process on a fine scale, it is hoped that these models can accurately learn the relationships between covariates and response at a fine spatial scale. However, validating these high resolution predictions can be a challenge, as often there is no data observed at this spatial scale. In this study, disaggregation regression was performed on simulated data in various settings and the resulting fine-scale predictions are compared to the simulated ground truth. Performance was investigated with varying numbers of data points, sizes of aggregated areas and levels of model misspecification. The effectiveness of cross validation on the aggregate level as a measure of fine-scale predictive performance was also investigated. Predictive performance improved as the number of observations increased and as the size of the aggregated areas decreased. When the model was well-specified, fine-scale predictions were accurate even with small numbers of observations and large aggregated areas. Under model misspecification predictive performance was significantly worse for large aggregated areas but remained high when response data was aggregated over smaller regions. Cross-validation correlation on the aggregate level was a moderately good predictor of fine-scale predictive performance. While these simulations are unlikely to capture the nuances of real-life response data, this study gives insight into the effectiveness of disaggregation regression in different contexts.