Enhancing Predictive Modeling of Nested Spatial Data through Group-Level Feature Disaggregation

Enhancing Predictive Modeling of Nested Spatial Data through Group-Level Feature Disaggregation
复制标题

DOI:
10.1145/3219819.3220091
复制
发表时间:
2018-07
期刊:
Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
影响因子:
--
通讯作者:
Boyang Liu;P. Tan;Jiayu Zhou
Boyang Liu;P. Tan;Jiayu Zhou
中科院分区:
其他
文献类型:
--
作者:
Boyang Liu;P. Tan;Jiayu Zhou

文献摘要

被引文献

相似文献

多层建模和多任务学习是两种广泛使用的嵌套(多层)数据建模方法,其中包含可以聚类成组的观察,其特征在于它们的组级特征。尽管他们解决的问题的相似性,多层次建模和多任务学习之间的明确关系还没有仔细检查。在本文中,我们提出了两种方法之间的比较分析,以说明它们的优势和局限性时,适用于两个层次的嵌套数据。我们提供了一个详细的分析,从优化的角度来看,在温和的条件下,证明他们的配方的等效性。我们还证明了它们在预测性能方面的局限性,特别是当应用于具有少量组或每组有限训练示例的数据集时,它们难以识别局部和组级特征之间的潜在跨尺度相互作用。为了克服这些限制,我们提出了一种新的方法分解嵌套数据中的组级特征的粗尺度值。在合成数据和真实数据上的实验结果表明,分解的组级特征可以显著提高模型的预测精度,并更有效地识别跨尺度交互作用。
Multilevel modeling and multi-task learning are two widely used approaches for modeling nested (multi-level) data, which contain observations that can be clustered into groups, characterized by their group-level features. Despite the similarity of the problems they address, the explicit relationship between multilevel modeling and multi-task learning has not been carefully examined. In this paper, we present a comparative analysis between the two methods to illustrate their strengths and limitations when applied to two-level nested data. We provide a detailed analysis demonstrating the equivalence of their formulations under a mild condition from an optimization perspective. We also demonstrate their limitations in terms of their predictive performance and especially, their difficulty in identifying potential cross-scale interactions between the local and group-level features when applied to datasets with either a small number of groups or limited training examples per group. To overcome these limitations, we propose a novel method for disaggregating the coarse-scale values of the group-level features in the nested data. Experimental results on both synthetic and real-world data show that the disaggregated group-level features can help enhance the prediction accuracy of the models significantly and identify the cross-scale interactions more effectively.