Best-effort semantic document search on GPUs

Best-effort semantic document search on GPUs
复制标题

GPU 上的尽力语义文档搜索

DOI:
10.1145/1735688.1735705
复制
发表时间:
2010
期刊:
Proceedings of the 2nd Workshop on Sustainable Computer Systems
影响因子:
--
通讯作者:
S. Cadambi
S. Cadambi
中科院分区:
--
文献类型:
--
作者:
S. Byna;Jiayuan Meng;A. Raghunathan;S. Chakradhar;S. Cadambi

文献摘要

被引文献

相似文献

语义索引是一种流行技术,用于访问和组织大量的非结构化文本数据。我们描述了在许多核GPU平台上的语义索引和文档搜索的优化实现。我们观察到,在128核Tesla C870 GPU上的语义索引的并行实现比在Intel Xeon 2.4GHz处理器上的顺序实现快2.4倍。在语义索引的工作负载特征和GPU的独特架构特征中,我们将不到壮观的速度归因于不匹配。与以很大的成功移植到GPU的常规数值计算相比,我们的语义索引算法(最近提议的监督语义索引算法称为SSI)具有有趣的特征 - 每个培训实例中的平行性量是数据依赖性和数据依赖性的,并且每次迭代都涉及一个具有稀疏向量的密集矩阵的乘积,从而导致随机内存访问模式。结果,我们观察到GPU实现的基线大大减少了GPU平台的硬件资源(处理元素和内存带宽)。但是,SSI算法也表现出独特的特征,我们将其共同称为算法的“宽容本质”。这些独特的特征允许新的优化,这些优化不努力将每个训练迭代的数值等效性与顺序实现保持相等。特别是,我们考虑了最佳及时计算技术,例如依赖性放松和计算下降,以适当地改变SSI的工作量特征,以利用GPU的独特建筑特征。我们还表明,在GPU上的依赖性放松和计算降低概念的实现与人们在多核心CPU上实现这些概念的方式完全不同,这在很大程度上是由于GPU支持的独特架构特征。我们的新技术极大地增加了平行工作量的数量,从而导致GPU的性能更高。通过优化CPU和GPU之间的数据传输,以及减少GPU内核调用开销,我们实现了进一步的性能增长。我们在Wikipedia超过180万个文档的数据库上评估了我们新的GPU加速搜索实施。通过应用我们的新型性能增强策略,我们对128核Tesla C870的GPU实施达到了5.5倍加速度,与对同一GPU的基线并行实施相比。与在双插头四核Intel Xeon多核CPU(8核)上的基线平行TBB实现相比,增强的GPU实现的速度更快为11倍。与同一多核CPU上的并行实现相比,该实现也使用数据依赖性放松和删除计算技术,我们增强的GPU实现速度更快。
Semantic indexing is a popular technique used to access and organize large amounts of unstructured text data. We describe an optimized implementation of semantic indexing and document search on manycore GPU platforms. We observed that a parallel implementation of semantic indexing on a 128-core Tesla C870 GPU is only 2.4X faster than a sequential implementation on an Intel Xeon 2.4GHz processor. We ascribe the less than spectacular speedup to a mismatch in the workload characteristics of semantic indexing and the unique architectural features of GPUs. Compared to the regular numerical computations that have been ported to GPUs with great success, our semantic indexing algorithm (the recently proposed Supervised Semantic Indexing algorithm called SSI) has interesting characteristics -- the amount of parallelism in each training instance is data-dependent, and each iteration involves the product of a dense matrix with a sparse vector, resulting in random memory access patterns. As a result, we observed that the baseline GPU implementation significantly under-utilizes the hardware resources (processing elements and memory bandwidth) of the GPU platform. However, the SSI algorithm also demonstrates unique characteristics, which we collectively refer to as the "forgiving nature" of the algorithm. These unique characteristics allow for novel optimizations that do not strive to preserve numerical equivalence of each training iteration with the sequential implementation. In particular, we consider best-effort computing techniques, such as dependency relaxation and computation dropping, to suitably alter the workload characteristics of SSI to leverage the unique architectural features of the GPU. We also show that the realization of dependency relaxation and computation dropping concepts on a GPU is quite different from how one would implement these concepts on a multicore CPU, largely due to the distinct architectural features supported by a GPU. Our new techniques dramatically enhance the amount of parallel workload, leading to much higher performance on the GPU. By optimizing data transfers between CPU and GPU, and by reducing GPU kernel invocation overheads, we achieve further performance gains. We evaluated our new GPU-accelerated implementation of semantic document search on a database of over 1.8 million documents from Wikipedia. By applying our novel performance-enhancing strategies, our GPU implementation on a 128-core Tesla C870 achieved a 5.5X acceleration as compared to a baseline parallel implementation on the same GPU. Compared to a baseline parallel TBB implementation on a dual-socket quad-core Intel Xeon multicore CPU (8-cores), the enhanced GPU implementation is 11X faster. Compared to a parallel implementation on the same multi-core CPU that also uses data dependency relaxation and dropping computation techniques, our enhanced GPU implementation is 5X faster.