Massive atomics for massive parallelism on GPUs

Massive atomics for massive parallelism on GPUs
复制标题

用于 GPU 上大规模并行性的大规模原子

DOI:
--
复制
发表时间:
2014
期刊:
International Symposium on Mathematical Morphology and Its Application to Signal and Image Processing
影响因子:
--
通讯作者:
E. Zhang
E. Zhang
中科院分区:
--
文献类型:
--
作者:
Ian J. Egielski;Jesse Huang;E. Zhang

文献摘要

被引文献

相似文献

许多应用中使用的一种重要的并行类型是归约型并行。在这些应用中,对一个共享数据对象的读取-修改-写入更新的顺序可以是任意的,只要存在读取-修改-写入更新的强加顺序即可。并行化这些类型的应用程序的典型方法是首先让每个单独的线程执行本地计算并将结果保存在线程私有数据对象中,然后在归约阶段合并来自所有工作线程的结果。所有适合 MapReduce 框架的应用程序都属于这一类。此外,机器学习、数据挖掘、数值分析和科学模拟应用程序也可能受益于约简型并行性。然而,通过使用线程私有数据对象的并行化方案在大规模并行 GPU 应用程序中可能不可行。由于并发线程数量极其庞大(至少数万个),线程私有数据对象创建可能会导致内存空间爆炸问题。 在本文中,我们提出了一种处理共享数据对象管理的新方法,以实现 GPU 上的简化型并行性。我们的方法利用细粒度并行性,同时保持良好的可编程性。它基于内部硬件原子指令的使用。原子操作可能看起来很昂贵,因为当多个线程同时原子地更新同一内存对象时,它会导致线程序列化。然而,我们发现,通过适当的原子碰撞减少技术,原子实现可以胜过非原子实现,即使对于已知具有高性能非原子 GPU 实现的基准测试也是如此。同时,原子的使用可以大大降低编码复杂性,因为线程私有对象管理或显式线程通信(对于受原子操作保护的共享数据对象)都不是必需的。
One important type of parallelism exploited in many applications is reduction type parallelism. In these applications, the order of the read-modify-write updates to one shared data object can be arbitrary as long as there is an imposed order for the read-modify-write updates. The typical way to parallelize these types of applications is to first let every individual thread perform local computation and save the results in thread-private data objects, and then merge the results from all worker threads in the reduction stage. All applications that fit into the map reduce framework belong to this category. Additionally, the machine learning, data mining, numerical analysis and scientific simulation applications may also benefit from reduction type parallelism. However, the parallelization scheme via the usage of thread-private data objects may not be vi- able in massively parallel GPU applications. Because the number of concurrent threads is extremely large (at least tens of thousands of), thread-private data object creation may lead to memory space explosion problems. In this paper, we propose a novel approach to deal with shared data object management for reduction type parallelism on GPUs. Our approach exploits fine-grained parallelism while at the same time maintaining good programmability. It is based on the usage of intrinsic hardware atomic instructions. Atomic operation may appear to be expensive since it causes thread serialization when multiple threads atomically update the same memory object at the same time. However, we discovered that, with appropriate atomic collision reduction techniques, the atomic implementation can out- perform the non-atomics implementation, even for benchmarks known to have high performance non-atomics GPU implementations. In the meantime, the usage of atomics can greatly reduce coding complexity as neither thread-private object management or explicit thread-communication (for the shared data objects protected by atomic operations) is necessary.