Efficiently mining maximal l-reachability co-location patterns from spatial data sets

Efficiently mining maximal l-reachability co-location patterns from spatial data sets
复制标题

DOI:
10.3233/ida-216515
复制
发表时间:
2023-01
期刊:
Intell. Data Anal.
影响因子:
--
通讯作者:
Muquan Zou;Lizhen Wang-;Pingping Wu;Vanha Tran
Muquan Zou;Lizhen Wang-;Pingping Wu;Vanha Tran
中科院分区:
其他
文献类型:
--
作者:
Muquan Zou;Lizhen Wang-;Pingping Wu;Vanha Tran

文献摘要

相似文献

协同定位模式是一组在空间上强相关的空间特征。然而,如果流行度量仅基于集团(或星星)关系,则这些模式中的一些可以被忽略。因此,通过引入l-可达团,提出了l-可达协同定位模式,其中每个实例对的成员可以在给定的步长l内彼此可达.由于l-可达性同位模式的平均长度趋于变长,本文研究了最大l-可达性同位模式挖掘。首先,一些稀疏化策略,以缩短星星邻居列表的实例在一个更新的图称为l-可达邻居关系图,然后,他们被分组,其相应的模式。其次,候选的最大l-可达性co-location模式迭代检测在一个大小无关的方式包含组密钥和它们的交集的bi-graphs。第三,每个候选最大l-可达性协同定位模式的流行度以二分搜索的方式用被称为1/2 l-可达性邻域列表的自然l-可达性集团来检查。最后,我们的模型和算法的有效性和效率进行了广泛的比较实验合成和真实世界的空间数据集。
A co-location pattern is a set of spatial features that are strongly correlated in space. However, some of these patterns could be neglected if the prevalence metrics are based solely on the clique (or star) relationship. Hence, the l-reachability co-location pattern is proposed by introducing the l-reachability clique where the members of each instance pair can be reachable to each other in a given step length l. Because the average size of l-reachability co-location patterns tends to be longer, maximal l-reachability co-location pattern mining is researched in this paper. First, some sparsification strategies are introduced to shorten star neighborhood lists of instances in an updated graph called the l-reachability neighbor relationship graph, and then, they are grouped by their corresponding patterns. Second, candidate maximal l-reachability co-location patterns are iteratively detected in a size-independent way on bi-graphs that contain group keys and their intersection sets. Third, the prevalence of each candidate maximal l-reachability co-location pattern is checked in a binary search way with a natural l-reachability clique called the ⌊l/2⌋-reachability neighborhood list. Finally, the effectiveness and efficiency of our model and algorithms are analyzed by extensive comparison experiments on synthetic and real-world spatial data sets.