Accelerating search and recognition workloads with SSE 4.2 string and text processing instructions

Accelerating search and recognition workloads with SSE 4.2 string and text processing instructions
复制标题

使用 SSE 4.2 字符串和文本处理指令加速搜索和识别工作负载

DOI:
10.1109/ispass.2011.5762731
复制
发表时间:
2011
期刊:
(IEEE ISPASS) IEEE INTERNATIONAL SYMPOSIUM ON PERFORMANCE ANALYSIS OF SYSTEMS AND SOFTWARE
影响因子:
--
通讯作者:
Mikko H. Lipasti
Mikko H. Lipasti
中科院分区:
--
文献类型:
--
作者:
Guangyu Shi;Min Li;Mikko H. Lipasti

文献摘要

被引文献

相似文献

今天的信息量增长迅速,每三年翻一番。因此,计算机应用程序中的搜索和识别阶段将消耗总CPU时间的越来越大的部分。SSE 4.2指令集首先在Intel的Core i7中实现,提供了字符串和文本处理指令(STTNI),这些指令利用SIMD操作来处理字符数据。虽然最初是为了加速字符串、文本和XML处理而设计的,但这些指令强大的新功能在这些领域之外也很有用,值得重新审视许多应用程序的搜索和识别阶段,以利用STTNI来提高性能。在本文中,我们探讨了使用STTNI来提高搜索和识别应用程序的CPU和内存性能的可行性和潜在好处。我们优化了四个基准应用程序-缓存模拟,B+树搜索算法,模板匹配,基本局部比对搜索工具(BLAST)-与STTNI,和新的应用程序优于其各自的原始实现的1.4倍至13倍。
Today's information is increasing rapidly, doubling every three years. Consequently, the search and recognition stages in computer applications will consume a growing portion of the total CPU time. The SSE 4.2 instruction set, first implemented in Intel's Core i7, provides string and text processing instructions (STTNI) that utilize SIMD operations for processing character data. Though originally conceived for accelerating string, text, and XML processing, the powerful new capabilities of these instructions are useful outside of these domains, and it is worth revisiting the search and recognition stages of numerous applications to utilize STTNI to improve performance. In this paper, we explored the feasibility and potential benefit of using STTNI to improve the CPU and memory performance of search-and-recognition applications. We optimized four benchmark applications — cache simulation, B+tree search algorithm, template matching, Basic Local Alignment Search Tool (BLAST) — with STTNI, and the new applications outperform their respective original implementations by a factor of 1.4× to 13×.