Multivariate Gaussian Random Number Generation Targeting Reconfigurable Hardware

Multivariate Gaussian Random Number Generation Targeting Reconfigurable Hardware
复制标题

DOI:
10.1145/1371579.1371584
复制
发表时间:
2008-06
期刊:
ACM Trans. Reconfigurable Technol. Syst.
影响因子:
--
通讯作者:
David B. Thomas;W. Luk
David B. Thomas;W. Luk
中科院分区:
其他
文献类型:
--
作者:
David B. Thomas;W. Luk

文献摘要

被引文献

相似文献

多变量高斯分布常被用来模拟随机时间序列之间的相关性,并可用于探索蒙特-卡罗模拟中N个时间序列之间的相关性的影响。然而,生成随机相关向量是一个O(N2)过程,并且很快成为软件模拟中的计算瓶颈。本文提出了一种在并行硬件中产生向量的有效方法,使用N个并行流水线组件每N个周期产生一个新向量。这种方法很好地映射到嵌入式块RAM和乘法器在当代FPGA中,特别是广泛的测试表明,有限的位宽算法不会降低所生成的矢量的统计质量。在Virtex-4架构中实现该架构实现了500 MHz的时钟速率,并且在最大的设备中可以支持高达512的向量长度。高时钟速率和并行性的结合提供了优于传统处理器的显着性能优势,500 MHz的xc 4vsx 55器件使用AMD优化的BLAS封装,提供了比Opteron 2.6GHz快200倍的加速。在Delta-Gamma Value-at Risk的案例研究中,使用xc 4vsx 55的400 MHz RC 2000加速卡比四路Opteron 2.6GHz SMP快26倍。
The multivariate Gaussian distribution is often used to model correlations between stochastic time-series, and can be used to explore the effect of these correlations across N time-series in Monte-Carlo simulations. However, generating random correlated vectors is an O(N2) process, and quickly becomes a computational bottleneck in software simulations. This article presents an efficient method for generating vectors in parallel hardware, using N parallel pipelined components to generate a new vector every N cycles. This method maps well to the embedded block RAMs and multipliers in contemporary FPGAs, particularly as extensive testing shows that the limited bit-width arithmetic does not reduce the statistical quality of the generated vectors. An implementation of the architecture in the Virtex-4 architecture achieves a 500MHz clock-rate, and can support vector lengths up to 512 in the largest devices. The combination of a high clock-rate and parallelism provides a significant performance advantage over conventional processors, with an xc4vsx55 device at 500MHz providing a 200 times speedup over an Opteron 2.6GHz using an AMD optimised BLAS package. In a case study in Delta-Gamma Value-at Risk, an RC2000 accelerator card using an xc4vsx55 at 400MHz is 26 times faster than a quad Opteron 2.6GHz SMP.