Statistically-Guided Deep Network Transformation to Harness Heterogeneity in Space (Extended Abstract)

Statistically-Guided Deep Network Transformation to Harness Heterogeneity in Space (Extended Abstract)
复制标题

DOI:
10.24963/ijcai.2022/752
复制
发表时间:
2022-07
期刊:
--
影响因子:
--
通讯作者:
Yiqun Xie;Erhu He;X. Jia;Han Bao;Xun Zhou;Rahul Ghosh;Praveen Ravirathinam
Yiqun Xie;Erhu He;X. Jia;Han Bao;Xun Zhou;Rahul Ghosh;Praveen Ravirathinam
中科院分区:
其他
文献类型:
--
作者:
Yiqun Xie;Erhu He;X. Jia;Han Bao;Xun Zhou;Rahul Ghosh;Praveen Ravirathinam

文献摘要

相似文献

空间数据无处不在,已经改变了许多关键领域的决策,包括公共卫生、农业、交通等。虽然机器学习的最新进展为利用海量空间数据集(例如卫星图像)提供了有希望的方法,但空间异质性--空间数据的一个基本特性--构成了一个重大挑战,因为数据分布或生成过程往往在空间上不同。最近针对这一困难问题的研究要么需要已知的空间划分作为输入,要么只能支持有限的特殊情况(例如,二进制分类)。此外,这些方法学习的异质性模式局限于训练样本的位置,不能应用于新的位置。我们提出了一个统计引导的框架,在训练过程中使用分布驱动的优化自适应地划分数据,并将深度学习模型(用户选择)转换为异构性感知的体系结构。我们还提出了一个空间调节器来将学习到的模式推广到新的测试区域。在真实数据集上的实验结果表明,该框架能够有效地捕捉异构性足迹,显著提高预测性能。
Spatial data are ubiquitous and have transformed decision-making in many critical domains, including public health, agriculture, transportation, etc. While recent advances in machine learning offer promising ways to harness massive spatial datasets (e.g., satellite imagery), spatial heterogeneity -- a fundamental property of spatial data -- poses a major challenge as data distributions or generative processes often vary over space. Recent studies targeting this difficult problem either require a known space-partitioning as the input, or can only support limited special cases (e.g., binary classification). Moreover, heterogeneity-pattern learned by these methods are locked to the locations of the training samples, and cannot be applied to new locations. We propose a statistically-guided framework to adaptively partition data in space during training using distribution-driven optimization and transform a deep learning model (of user's choice) into a heterogeneity-aware architecture. We also propose a spatial moderator to generalize learned patterns to new test regions. Experiment results on real-world datasets show that the framework can effectively capture footprints of heterogeneity and substantially improve prediction performances.