Statistical Inference for Spatial Regionalization

Statistical Inference for Spatial Regionalization
复制标题

DOI:
10.1145/3589132.3625608
复制
发表时间:
2023-11
期刊:
Proceedings of the 31st ACM International Conference on Advances in Geographic Information Systems
影响因子:
--
通讯作者:
Hussah Alrashid;A. Magdy;Sergio Rey
Hussah Alrashid;A. Magdy;Sergio Rey
中科院分区:
其他
文献类型:
--
作者:
Hussah Alrashid;A. Magdy;Sergio Rey

文献摘要

相似文献

区域化的过程包括将一组空间区域聚集成空间上连续的区域。鉴于区域化问题的NP-Hard性质,所有现有的算法都给出了近似解。为了确定这些近似的质量,与从所有潜在样本解得出的随机参考分布相比,领域专家获得关于优化目标函数的统计上有意义的证据是至关重要的。本文提出了一种新的空间区域化问题,称为空间区域化统计推断(SISR),它产生具有预定区域基数的随机样本解。SISR背后的驱动力是对任何给定的区域化方案进行统计推断。为了解决SISR问题,我们提出了一种称为PRRP的并行技术(P-区域化通过递归划分)。PRRP分三个阶段运行:区域生长阶段构建具有预定义基数的初始区域,而区域合并和区域分割阶段确保未分配区域的空间连续性,允许具有预定义基数的后续区域的增长。一项广泛的评估显示了使用各种真实数据集的PRRP的有效性。
The process of regionalization involves clustering a set of spatial areas into spatially contiguous regions. Given the NP-hard nature of regionalization problems, all existing algorithms yield approximate solutions. To ascertain the quality of these approximations, it is crucial for domain experts to obtain statistically significant evidence on optimizing the objective function, in comparison to a random reference distribution derived from all potential sample solutions. In this paper, we propose a novel spatial regionalization problem, denoted as SISR (Statistical Inference for Spatial Regionalization), which generates random sample solutions with a predetermined region cardinality. The driving motivation behind SISR is to conduct statistical inference on any given regionalization scheme. To address SISR, we present a parallel technique named PRRP (P-Regionalization through Recursive Partitioning). PRRP operates over three phases: the region growing phase constructs initial regions with a predefined cardinality, while the region merging and region splitting phases ensure the spatial contiguity of unassigned areas, allowing for the growth of subsequent regions with predefined cardinalites. An extensive evaluation shows the effectiveness of PRRP using various real datasets.