课题基金 / 基金详情

SHF:Small: Solving the Problem of Scalable Multi-Precision Matrix Arithmetic on GPUs

SHF:Small: Solving the Problem of Scalable Multi-Precision Matrix Arithmetic on GPUs
SHF:Small:解决 GPU 上可扩展多精度矩阵算术问题
批准号:
1217590
负责人:
Charles Weems
金额:
$45.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-06-01 至 2016-05-31

项目摘要

项目成果

Charles Weems的其他基金

相似基金

相关文献

中文摘要
翻译
计算机直接支持通常仅限于位(约19位小数位)精度的算术。需要更高精度的应用程序必须通过计算成本高昂的软件来实现算术。超过大约256位的精度,这样的计算变得相当昂贵。例如,RSA加密算法可能需要高达4096位精度的算术。在实验数学和数论等领域的应用可能需要数百万位的精度。在现代处理器上,一次1000万位精度的乘法可能需要十分之一秒的计算时间,这意味着使用如此大的值的矩阵运算可能需要几天到几周的时间才能执行。在以前的工作中,研究人员已经证明,通过利用商用图形处理单元(GPU)的并行处理能力来取代传统的CPU,可以将性能提高20倍。然而,对GPU进行编程以实现这种级别的性能是相当困难的,生成的代码需要大量手动调整才能将其移植到新一代GPU并获得其性能优势,其扩展速度正在超过CPU性能扩展。该项目正在努力开发一个框架,该框架可以自动生成和调整多精度算术库,以便在连续几代GPU上执行。这些库包括标量和基本矩阵算术例程。它们支持在精度和矩阵大小方面进行缩放。这个问题具有挑战性,因为必须为不同的精度水平自动选择不同的并行算法,这必须与开发矩阵运算固有的并行性的交替维度相平衡。此外,这项工作寻求在使用GPU增强的计算机集群中采用分布式并行,以便这些库可以用于开始在国家实验室部署的基于GPU的新一代超级计算机上。这项工作意义重大,因为它使开发低成本商用图形处理器变得更容易,从而使多精度标量和矩阵运算的性能提高了一个数量级以上。一个重要的应用是增强RSA加密的性能,以更高的数据速率支持更长、更安全的密钥,从而使加密更大量的互联网流量变得可行。另一个重要用途是实验数学,在实验数学中,计算代价高昂的函数(例如,积分、无穷级数)以高精度计算,并与其他函数和高精度常量进行比较,以帮助确定更有效的闭合形式解。实验数学的结果在粒子物理、混沌理论和基本常数的计算中得到了应用。由此产生的软件框架为从个人研究人员工作站到大型超级计算机的各种系统提供了显著的多精度算术性能增强。
英文摘要
Computers directly support arithmetic that is typically limited to 64 bits (about 19 decimal digits) of precision. Applications that need more precision must implement arithmetic through computationally expensive software. Beyond about 256 bits of precision, such calculations become quite costly. The RSA encryption algorithm, for example, can require arithmetic with up to 4096 bits of precision. Applications in areas such as experimental mathematics and number theory can require millions of bits of precision. One multiplication with 10 million bits of precision can take a tenth of a second to compute on a modern processor, which means that matrix arithmetic using such large values can take days to weeks to execute. In previous work the investigators have shown that it is possible to obtain a factor of 20 improvement in performance by utilizing the parallel processing capabilities of a commodity graphics processing unit (GPU) in place of the traditional CPU. However, programming a GPU to achieve this level of performance is quite difficult, and the resulting code requires considerable hand-tuning to move it to new generations of GPU and gain the advantage of their performance, which is scaling up at a rate that exceeds CPU performance scaling.This project is working to develop a framework that automatically generates and tunes multi-precision arithmetic libraries to execute on successive generations of GPUs. The libraries include both scalar and basic matrix arithmetic routines. They support scaling in precision as well as matrix size. The problem is challenging because different parallel algorithms must be automatically selected for different levels of precision, which must be balanced with the exploitation of the alternate dimension of parallelism inherent in matrix arithmetic. In addition, the work seeks to employ distributed parallelism across a cluster of computers enhanced with GPUs, so that the libraries can be used on a new generation of GPU-based supercomputers that is beginning to be deployed at national laboratories. The work is significant because it enables easier exploitation of low-cost commodity graphics processors to achieve more than an order of magnitude increase in performance for multi-precision scalar and matrix arithmetic. One important application is enhancing performance of RSA encryption to support longer, more secure keys, at greater data rates, so that it becomes feasible to encrypt greater volumes of internet traffic. Another important use is experimental mathematics, where computationally expensive functions (e.g., integrals, infinite series) are computed at high precision and compared to other functions and high precision constants to help identify more efficient closed-form solutions. Results from experimental mathematics have found applications in particle physics, chaos theory, and calculation of fundamental constants. The resulting software framework offers a significant performance enhancement for multi-precision arithmetic to systems that range from individual researcher workstations to large supercomputers.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research:CyberTraining:Implementation:Medium: Modern Course Exemplars infused with Parallel and Distributed Computing for the Introductory Computing Course Sequence
  • 批准号:
    2321016
  • 项目类别:
    Standard Grant
  • 资助金额:
    $17.8万
  • 财政年份:
    2023
  • 负责人:
    Charles Weems
  • 依托单位:
Collaborative Research:CyberTraining: Implementation: Medium:Broadening Adoption of Parallel and Distributed Computing in Undergraduate Computer Science and Engineering Curricula
  • 批准号:
    2017427
  • 项目类别:
    Standard Grant
  • 资助金额:
    $16.0万
  • 财政年份:
    2020
  • 负责人:
    Charles Weems
  • 依托单位:
Collaborative Research:CyberTraining:Conceptualization: Planning a Sustainable Ecosystem for Incorporating Parallel and Distributed Computing into Undergraduate Education
  • 批准号:
    1924023
  • 项目类别:
    Standard Grant
  • 资助金额:
    $3.18万
  • 财政年份:
    2019
  • 负责人:
    Charles Weems
  • 依托单位:
Collaborative Research: CyberTraining: CDL: Preparing Instructors to Offer Experimental Courses in an Updated PDC Curriculum, and Broadening Participation
  • 批准号:
    1730527
  • 项目类别:
    Standard Grant
  • 资助金额:
    $36.73万
  • 财政年份:
    2017
  • 负责人:
    Charles Weems
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: