Bit-Parallel Approximate Pattern Matching on the Xeon Phi Coprocessor

Bit-Parallel Approximate Pattern Matching on the Xeon Phi Coprocessor
复制标题

Xeon Phi 协处理器上的位并行近似模式匹配

DOI:
10.1109/sbac-pad.2014.37
复制
发表时间:
2014
期刊:
2014 IEEE 26th International Symposium on Computer Architecture and High Performance Computing
影响因子:
--
通讯作者:
B. Schmidt
B. Schmidt
中科院分区:
--
文献类型:
--
作者:
T. T. Tran;Simon Schindel;Yongchao Liu;B. Schmidt

文献摘要

被引文献

相似文献

位并行模式匹配在位数组中编码计算值。这种方法通过在一个机器字内执行多次更新来提高效率。因此,一个重要的参数是机器字大小(例如32或64位)。随着向量寄存器长度的增加,位并行模式匹配算法到现代高性能计算架构的有效映射变得越来越重要。本文研究了Wu-Manber近似模式匹配算法在Intel Xeon Phi协处理器上的高效实现。该架构具有512位长的向量处理单元(VPU)以及大量处理核心。我们提出了两个映射的Wu-Manber算法的基础上自动向量化和intrinsic,分别。我们的评估表明,内在的方法产生更高的性能,并获得约两个数量级的串行CPU版本相比,加速。源代码可在http://xbitpar.sourceforge.net/上获得。
Bit-parallel pattern matching encodes calculated values in bit arrays. This approach gains its efficiency by performing multiple updates within a machine word. An important parameter is therefore the machine word size (e.g. 32 or 64 bits). With the increasing length of vector registers, the efficient mapping of bit-parallel pattern matching algorithms onto modern high performance computing architectures is becoming increasingly important. In this paper, we investigate an efficient implementation of the Wu-Manber approximate pattern matching algorithm on the Intel Xeon Phi coprocessor. This architecture features a 512-bit long vector processing unit (VPU) as well as a large number of processing cores. We present two mappings of the Wu-Manber algorithm based on auto-vectorization and intrinsics, respectively. Our evaluation shows that the intrinsic approach yields higher performance and gains a speedup of around two orders-of-magnitude compared to a serial CPU version. The source code is available at http://xbitpar.sourceforge.net/.