ALICE HLT High Speed Tracking on GPU
ALICE HLT High Speed Tracking on GPU
复制标题
ALICE HLT GPU 上的高速跟踪
DOI:
--
复制
发表时间:
2011
影响因子:
1.8
通讯作者:
Alice Collaboration
中科院分区:
文献类型:
--
作者:
S. Gorbunov;David Rohr;K. Aamodt;T. Alt;H. Appelshäuser;A. Arend;M. Bach;Bruce Becker;S. Böttger;T. Breitner;H. Büsching;S. Chattopadhyay;J. Cleymans;C. Cicalò;I. Das;Ø. Djuvsland;H. Engel;H. Erdal;R. Fearick;Ø. Haaland;P. Hille;S. Kalcher;K. Kanaki;U. Kebschull;I. Kisel;M. Kretz;C. Lara;S. Lindal;V. Lindenstruth;A. A. Masoodi;G. Øvrebekk;R. Panse;Jörg Peschek;M. Płoskoń;T. Pocheptsov;Dinesh Ram;Theodor Rascanu;M. Richter;D. Röhrich;F. Ronchetti;B. Skaali;Olav Smørholm;C. Stokkevåg;T. Steinbeck;A. Szostak;J. Thäder;T. Tveter;K. Ullaland;Z. Vilakazi;Robert Weis;Z. Yin;P. Zelnicek;Alice Collaboration
The on-line event reconstruction in ALICE is performed by the High Level Trigger, which should process up to 2000 events per second in proton-proton collisions and up to 300 central events per second in heavy-ion collisions, corresponding to an input data stream of 30 GB/s. In order to fulfill the time requirements, a fast on-line tracker has been developed. The algorithm combines a Cellular Automaton method being used for a fast pattern recognition and the Kalman Filter method for fitting of found trajectories and for the final track selection. The tracker was adapted to run on Graphics Processing Units (GPU) using the NVIDIA Compute Unified Device Architecture (CUDA) framework. The implementation of the algorithm had to be adjusted at many points to allow for an efficient usage of the graphics cards. In particular, achieving a good overall workload for many processor cores, efficient transfer to and from the GPU, as well as optimized utilization of the different memories the GPU offers turned out to be critical. To cope with these problems a dynamic scheduler was introduced, which redistributes the workload among the processor cores. Additionally a pipeline was implemented so that the tracking on the GPU, the initialization and the output processed by the CPU, as well as the DMA transfer can overlap. The GPU tracking algorithm significantly outperforms the CPU version for large events while it entirely maintains its efficiency.