High Influencing Pattern Discovery over Time Series Data

High Influencing Pattern Discovery over Time Series Data
复制标题

DOI:
10.3390/ijgi10100696
复制
发表时间:
2021-10
期刊:
ISPRS Int. J. Geo Inf.
影响因子:
--
通讯作者:
Dianwu Fang;Lizhen Wang;Jialong Wang;Meijiao Wang
Dianwu Fang;Lizhen Wang;Jialong Wang;Meijiao Wang
中科院分区:
其他
文献类型:
--
作者:
Dianwu Fang;Lizhen Wang;Jialong Wang;Meijiao Wang

文献摘要

相似文献

空间协同定位模式表示其实例频繁出现在附近的空间特征的子集。高影响力同位模式挖掘是用来发现在特定方面有较高影响力的同位模式。此类模式挖掘的研究通常依赖于空间距离来度量实例之间的贴近度,这种方法不能应用于从疫情传播情景得出的影响传播过程。为了利用该领域的丰硕成果来发现有意义的模式,我们扩展了现有的方法,并提出了一个挖掘框架。我们首先定义了贴近度的新概念来刻画不同特征实例之间的语义贴近度,从而应用星形物化模型来挖掘影响模式。然后,设计属性描述子,从时间序列数据中感知实例和边的属性,并通过层次分析法计算属性权重,从而计算实例之间的影响以及影响模式中特征的影响。接下来,我们构建了影响度量,并设置了一个阈值来发现高影响模式。针对度量不满足向下闭包性质的问题,本文提出了两种改进算法来提高效率。在真实数据集和合成数据集上进行的大量实验验证了该方法的有效性、高效性和可扩展性。
A spatial co-location pattern denotes a subset of spatial features whose instances frequently appear nearby. High influence co-location pattern mining is used to find co-location patterns with high influence in specific aspects. Studies of such pattern mining usually rely on spatial distance for measuring nearness between instances, a method that cannot be applied to an influence propagation process concluded from epidemic dispersal scenarios. To discover meaningful patterns by using fruitful results in this field, we extend existing approaches and propose a mining framework. We first defined a new concept of proximity to depict semantic nearness between instances of distinct features, thus applying a star-shaped materialized model to mine influencing patterns. Then, we designed attribute descriptors to perceive attributes of instances and edges from time series data, and we calculated the attribute weights via an analytic hierarchy process, thereby computing the influence between instances and the influence of features in influencing patterns. Next, we constructed influencing metrics and set a threshold to discover high influencing patterns. Since the metrics do not satisfy the downward closure property, we propose two improved algorithms to boost efficiency. Extensive experiments conducted on real and synthetic datasets verified the effectiveness, efficiency, and scalability of our method.