A distributed multi-GPU system for high speed electron microscopic tomographic reconstruction.

A distributed multi-GPU system for high speed electron microscopic tomographic reconstruction.
复制标题

DOI:
10.1016/j.ultramic.2011.03.015
复制
发表时间:
2011-07
期刊:
影响因子:
2.2
通讯作者:
Agard, David A.
Agard, David A.
中科院分区:
工程技术3区
文献类型:
--
作者:
Zheng, Shawn Q.;Branlund, Eric;Kesthelyi, Bettina;Braunfeld, Michael B.;Cheng, Yifan;Sedat, John W.;Agard, David A.

文献摘要

参考文献

被引文献

相似文献

大尺度倾斜序列的全分辨率电子显微层析成像(EMT)重建需要强大的计算能力。执行迭代重建和重新排列的多个周期的愿望极大地增加了改进重建性能的迫切需要。这促使我们开发一种分布式多gpu(图形处理单元)系统,为非常大的三维(3D)体积的快速受限、迭代重建提供所需的计算能力。参与的gpu并行重构体的片段,然后将片段组装成完整的三维体。由于其强大的功能和多功能性,我们选择了CUDA (NVIDIA, USA)平台作为EMT重建的GPU实现。对于包含5张GTX295卡提供的10个gpu的系统,从包含122张40962像素投影图像(单精度浮点)的输入倾斜序列中获得40962 × 512体素的断层图,10个周期的SIRT重建总共需要1845秒,其中1032秒用于计算,其余为系统开销。同样的系统只需要39秒就可以从122 10242像素的投影中重建10242 × 256体素。虽然系统开销不是微不足道的,但性能分析表明,向系统添加额外的gpu将导致整体性能的稳步提高。因此,该系统可以很容易地扩展,为非常大的层析重建产生优越的计算能力,特别是赋予重建和重新排列的迭代周期。
Full resolution electron microscopic tomographic (EMT) reconstruction of large-scale tilt series requires significant computing power. The desire to perform multiple cycles of iterative reconstruction and realignment dramatically increases the pressing need to improve reconstruction performance. This has motivated us to develop a distributed multi-GPU (graphics processing unit) system to provide the required computing power for rapid constrained, iterative reconstructions of very large three-dimensional (3D) volumes. The participating GPUs reconstruct segments of the volume in parallel, and subsequently, the segments are assembled to form the complete 3D volume. Owing to its power and versatility, the CUDA (NVIDIA, USA) platform was selected for GPU implementation of the EMT reconstruction. For a system containing 10 GPUs provided by 5 GTX295 cards, 10 cycles of SIRT reconstruction for a tomogram of 40962 × 512 voxels from an input tilt series containing 122 projection images of 40962 pixels (single precision float) takes a total of 1845 seconds of which 1032 seconds are for computation with the remainder being the system overhead. The same system takes only 39 seconds total to reconstruct 10242 × 256 voxels from 122 10242 pixel projections. While the system overhead is non-trivial, performance analysis indicates that adding extra GPUs to the system would lead to steadily enhanced overall performance. Therefore, this system can be easily expanded to generate superior computing power for very large tomographic reconstructions and especially to empower iterative cycles of reconstruction and realignment.
DOI: 10.1016/j.jsb.2009.03.019
发表时间: 2009-07
影响因子: 3
作者:
Suloway, Christian;Shi, Jian;Cheng, Anchi;Pulokas, James;Carragher, Bridget;Potter, Clinton S.;Zheng, Shawn Q.;Agard, David A.;Jensen, Grant J.
通讯作者: Jensen, Grant J.
DOI: 10.1016/j.jsb.2009.06.010
发表时间: 2009-11-01
影响因子: 3
作者:
Zheng, Shawn Q.;Matsuda, Atsushi;Agard, David A.
通讯作者: Agard, David A.
DOI: 10.1016/0161-7346(84)90008-7
发表时间: 1984-01-01
期刊: ULTRASONIC IMAGING
影响因子: 2.3
作者:
ANDERSEN, AH;KAK, AC
通讯作者: KAK, AC
DOI: 10.1016/j.jsb.2010.01.008
发表时间: 2010-06-01
影响因子: 3
作者:
Agulleiro, J. I.;Garzon, E. M.;Fernandez, J. J.
通讯作者: Fernandez, J. J.
DOI: 10.1016/j.jsb.2006.08.010
发表时间: 2007-01-01
影响因子: 3
作者:
Diez, Daniel Castano;Mueller, Hannes;Frangakis, Achilleas S.
通讯作者: Frangakis, Achilleas S.