High-Performance Statistical Computing in the Computing Environments of the 2020s.

High-Performance Statistical Computing in the Computing Environments of the 2020s.
复制标题

DOI:
10.1214/21-sts835
复制
发表时间:
2022-11
期刊:
Statistical science : a review journal of the Institute of Mathematical Statistics
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

过去十年的技术进步,无论是硬件还是软件,都使得访问高性能计算(HPC)比以往任何时候都更容易。我们从统计计算的角度来回顾这些进展。云计算使超级计算机的使用变得负担得起。深度学习软件库使编程统计算法变得简单,并使用户能够编写一次代码并在任何地方运行-从笔记本电脑到具有多个图形处理单元(GPU)的工作站或云中的超级计算机。突出这些发展如何有利于统计学家,我们回顾了最近的优化算法,是有用的高维模型,可以利用HPC的力量。提供代码片段来演示编程的简易性。我们还提供了一个易于使用的分布式矩阵数据结构,适用于HPC。采用这种数据结构,我们说明了各种统计应用,包括大规模的正电子发射断层扫描和E1-正则化考克斯回归。我们的示例可以轻松扩展到云中的8 GPU工作站和720 CPU核心集群。作为一个恰当的例子,我们分析了2型糖尿病的发病从英国生物银行与200,000名受试者和约500,000单核苷酸多态性,使用HPC E31-正则化考克斯回归。拟合这个50万变量的模型需要不到45分钟的时间,并重新确认已知的关联。据我们所知,这是第一次证明在这个尺度上惩罚性回归生存结局的可行性。
Technological advances in the past decade, hardware and software alike, have made access to high-performance computing (HPC) easier than ever. We review these advances from a statistical computing perspective. Cloud computing makes access to supercomputers affordable. Deep learning software libraries make programming statistical algorithms easy and enable users to write code once and run it anywhere—from a laptop to a workstation with multiple graphics processing units (GPUs) or a supercomputer in a cloud. Highlighting how these developments benefit statisticians, we review recent optimization algorithms that are useful for high-dimensional models and can harness the power of HPC. Code snippets are provided to demonstrate the ease of programming. We also provide an easy-to-use distributed matrix data structure suitable for HPC. Employing this data structure, we illustrate various statistical applications including large-scale positron emission tomography and ℓ1-regularized Cox regression. Our examples easily scale up to an 8-GPU workstation and a 720-CPU-core cluster in a cloud. As a case in point, we analyze the onset of type-2 diabetes from the UK Biobank with 200,000 subjects and about 500,000 single nucleotide polymorphisms using the HPC ℓ1-regularized Cox regression. Fitting this half-million-variate model takes less than 45 minutes and reconfirms known associations. To our knowledge, this is the first demonstration of the feasibility of penalized regression of survival outcomes at this scale.
DOI: 10.1007/s10107-013-0697-1
发表时间: 2014-08-01
影响因子: 2.7
作者:
Chi, Eric C.;Zhou, Hua;Lange, Kenneth
通讯作者: Lange, Kenneth
DOI: 10.1007/s10444-011-9254-8
发表时间: 2013-04-01
影响因子: 1.7
作者:
Bang Cong Vu
通讯作者: Bang Cong Vu
DOI: 10.1137/080716542
发表时间: 2009-01-01
影响因子: 2.1
作者:
Beck, Amir;Teboulle, Marc
通讯作者: Teboulle, Marc
DOI: 10.1109/tpds.2018.2872064
发表时间: 2019-04-01
影响因子: 5.3
作者:
Besard, Tim;Foket, Christophe;De Sutter, Bjorn
通讯作者: De Sutter, Bjorn
DOI: 10.1093/bioinformatics/btp608
发表时间: 2010-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Buckner, Joshua;Wilson, Justin;Meng, Fan
通讯作者: Meng, Fan