ALICE HLT High Speed Tracking on GPU

ALICE HLT High Speed Tracking on GPU
复制标题

ALICE HLT GPU 上的高速跟踪

DOI:
--
复制
发表时间:
2011
影响因子:
1.8
通讯作者:
Alice Collaboration
Alice Collaboration
中科院分区:
工程技术3区
文献类型:
--
作者:
S. Gorbunov;David Rohr;K. Aamodt;T. Alt;H. Appelshäuser;A. Arend;M. Bach;Bruce Becker;S. Böttger;T. Breitner;H. Büsching;S. Chattopadhyay;J. Cleymans;C. Cicalò;I. Das;Ø. Djuvsland;H. Engel;H. Erdal;R. Fearick;Ø. Haaland;P. Hille;S. Kalcher;K. Kanaki;U. Kebschull;I. Kisel;M. Kretz;C. Lara;S. Lindal;V. Lindenstruth;A. A. Masoodi;G. Øvrebekk;R. Panse;Jörg Peschek;M. Płoskoń;T. Pocheptsov;Dinesh Ram;Theodor Rascanu;M. Richter;D. Röhrich;F. Ronchetti;B. Skaali;Olav Smørholm;C. Stokkevåg;T. Steinbeck;A. Szostak;J. Thäder;T. Tveter;K. Ullaland;Z. Vilakazi;Robert Weis;Z. Yin;P. Zelnicek;Alice Collaboration

文献摘要

被引文献

相似文献

爱丽丝中的在线事件重建是由高级触发器执行的,该触发器应在质子 - 普罗顿碰撞中每秒处理高达2000个事件,在重合离子碰撞中每秒最多300个中心事件,与输入数据流相对应30 GB/s。为了满足时间要求,已经开发了快速的在线跟踪器。该算法结合了一种用于快速模式识别的细胞自动机方法和用于拟合发现的轨迹和最终轨道选择的Kalman滤波器方法。该跟踪器适用于使用NVIDIA Compute Unified设备体系结构(CUDA)框架在图形处理单元(GPU)上运行。必须在许多方面调整算法的实现,以便有效使用图形卡。特别是,为许多处理器内核,有效地转移到GPU,以及对GPU提供的不同记忆的优化利用非常重要。为了解决这些问题,引入了动态调度程序,该调度程序将处理器内核之间的工作量重新分配。此外,还实施了管道,以便在GPU上进行跟踪,CPU处理的初始化和输出以及DMA传输可能会重叠。 GPU跟踪算法的表现明显胜过大型事件的CPU版本,同时它完全保持其效率。
The on-line event reconstruction in ALICE is performed by the High Level Trigger, which should process up to 2000 events per second in proton-proton collisions and up to 300 central events per second in heavy-ion collisions, corresponding to an input data stream of 30 GB/s. In order to fulfill the time requirements, a fast on-line tracker has been developed. The algorithm combines a Cellular Automaton method being used for a fast pattern recognition and the Kalman Filter method for fitting of found trajectories and for the final track selection. The tracker was adapted to run on Graphics Processing Units (GPU) using the NVIDIA Compute Unified Device Architecture (CUDA) framework. The implementation of the algorithm had to be adjusted at many points to allow for an efficient usage of the graphics cards. In particular, achieving a good overall workload for many processor cores, efficient transfer to and from the GPU, as well as optimized utilization of the different memories the GPU offers turned out to be critical. To cope with these problems a dynamic scheduler was introduced, which redistributes the workload among the processor cores. Additionally a pipeline was implemented so that the tracking on the GPU, the initialization and the output processed by the CPU, as well as the DMA transfer can overlap. The GPU tracking algorithm significantly outperforms the CPU version for large events while it entirely maintains its efficiency.