Parallel Sequence Mining on Shared-Memory Machines

Parallel Sequence Mining on Shared-Memory Machines
复制标题

DOI:
10.1006/jpdc.2000.1695
复制
发表时间:
1999-08
期刊:
J. Parallel Distributed Comput.
影响因子:
--
通讯作者:
Mohammed J. Zaki
Mohammed J. Zaki
中科院分区:
其他
文献类型:
--
作者:
Mohammed J. Zaki

文献摘要

被引文献

相似文献

我们提出了pSPADE,在大型数据库中快速发现频繁序列的并行算法。pSPADE将原始搜索空间分解为更小的基于后缀的类。每个类都可以在主内存中使用高效的搜索技术和简单的连接操作来解决。此外,每个类可以在每个处理器上独立求解,无需同步。但是,必须利用动态类间和类内负载平衡来确保每个处理器获得等量的工作。在一个12处理器的SGI Origin 2000共享内存系统上的实验表明,该算法具有良好的加速比和扩展性能。
We present pSPADE, a parallel algorithm for fast discovery of frequent sequences in large databases. pSPADE decomposes the original search space into smaller suffix-based classes. Each class can be solved in main-memory using efficient search techniques, and simple join operations. Further each class can be solved independently on each processor requiring no synchronization. However, dynamic inter-class and intra-class load balancing must be exploited to ensure that each processor gets an equal amount of work. Experiments on a 12 processor SGI Origin 2000 shared memory system show good speedup and excellent scaleup results.