Spatial Ensemble Learning for Heterogeneous Geographic Data with Class Ambiguity

Spatial Ensemble Learning for Heterogeneous Geographic Data with Class Ambiguity
复制标题

DOI:
10.1145/3337798
复制
发表时间:
2019-08-01
影响因子:
5
通讯作者:
Knight, Joseph
Knight, Joseph
中科院分区:
计算机科学3区
文献类型:
--
作者:
Jiang, Zhe;Sainju, Arpan Man;Knight, Joseph

文献摘要

被引文献

相似文献

类歧义是指类似特征对应于不同位置的不同类别的现象。鉴于具有阶级歧义的异质地理数据,空间集合学习(SEL)问题旨在将地理区域分解为分离区域,以使阶级模棱两可最小化,并且可以在每个区域中学习局部分类器。这个问题对于诸如从异质地球观察数据中映射的土地覆盖物诸如频谱混乱的应用很重要。但是,由于其较高的计算成本,问题具有挑战性。合奏学习中的相关工作要么假定相同的样本分布(例如,包装,增强,随机森林)或分解特征矢量空间中的多模数输入数据(例如,专家的混合物,多模式集合),因此无法有效地最大程度地减少班级歧义。相比之下,我们提出了一个空间集合框架,该框架在地理空间中明确划分了输入数据。我们的方法首先将数据预处理成均匀的空间斑块,并使用贪婪的启发式方法将具有高级歧义的斑块分配给不同区域。我们进一步扩展了基于空间自相关效应的附近区域之间的空间集合学习框架,并在附近区域之间进行空间依赖。对两个现实世界湿地映射数据集进行的理论分析和实验评估都表明该方法的可行性。
Class ambiguity refers to the phenomenon whereby similar features correspond to different classes at different locations. Given heterogeneous geographic data with class ambiguity, the spatial ensemble learning (SEL) problem aims to find a decomposition of the geographic area into disjoint zones such that class ambiguity is minimized and a local classifier can be learned in each zone. The problem is important for applications such as land cover mapping from heterogeneous earth observation data with spectral confusion. However, the problem is challenging due to its high computational cost. Related work in ensemble learning either assumes an identical sample distribution (e.g., bagging, boosting, random forest) or decomposes multi-modular input data in the feature vector space (e.g., mixture of experts, multimodal ensemble) and thus cannot effectively minimize class ambiguity. In contrast, we propose a spatial ensemble framework that explicitly partitions input data in geographic space. Our approach first preprocesses data into homogeneous spatial patches and uses a greedy heuristic to allocate pairs of patches with high class ambiguity into different zones. We further extend our spatial ensemble learning framework with spatial dependency between nearby zones based on the spatial autocorrelation effect. Both theoretical analysis and experimental evaluations on two real world wetland mapping datasets show the feasibility of the proposed approach.