Mining Statistically Significant Co-location and Segregation Patterns

Mining Statistically Significant Co-location and Segregation Patterns
复制标题

DOI:
10.1109/tkde.2013.88
复制
发表时间:
2014-05
影响因子:
8.9
通讯作者:
Sajib Barua;J. Sander
Sajib Barua;J. Sander
中科院分区:
计算机科学2区
文献类型:
--
作者:
Sajib Barua;J. Sander

文献摘要

被引文献

相似文献

在空间域中,特征之间的相互作用产生两种类型的相互作用模式:共定位模式和分离模式。现有的寻找同址模式的方法有几个缺点:(1)它们依赖于用户指定的患病率测量阈值;(2)未考虑空间自相关;(3)即使特征是随机分布的,它们也可能报告共存。种族隔离模式尚未受到太多关注。在本文中,我们提出了一种基于统计检验的方法来发现这两种类型的交互模式。本文引入了一种新的共定位和分离模式的定义,提出了一种考虑空间自相关的特征零分布模型,并设计了一种同时发现共定位和分离模式的算法。与基于数据分布模拟的naïve方法相比,我们还开发了两种策略来降低计算成本,并且我们提出了一种方法,通过使用特征邻域的近似值来进一步减少算法的运行时间。我们使用合成和真实数据集对我们的方法进行了经验评估,并展示了其优于最先进的协同位置挖掘算法的优势。
In spatial domains, interaction between features gives rise to two types of interaction patterns: co-location and segregation patterns. Existing approaches to finding co-location patterns have several shortcomings: (1) They depend on user specified thresholds for prevalence measures; (2) they do not take spatial auto-correlation into account; and (3) they may report co-locations even if the features are randomly distributed. Segregation patterns have yet to receive much attention. In this paper, we propose a method for finding both types of interaction patterns, based on a statistical test. We introduce a new definition of co-location and segregation pattern, we propose a model for the null distribution of features so spatial auto-correlation is taken into account, and we design an algorithm for finding both co-location and segregation patterns. We also develop two strategies to reduce the computational cost compared to a naïve approach based on simulations of the data distribution, and we propose an approach to reduce the runtime of our algorithm even further by using an approximation of the neighborhood of features. We evaluate our method empirically using synthetic and real data sets and demonstrate its advantages over a state-of-the-art co-location mining algorithm.