课题基金 / 基金详情

ParaSol: Fine-Grained Thread-Level Parallelism for Single-Threaded Performance

ParaSol: Fine-Grained Thread-Level Parallelism for Single-Threaded Performance
ParaSol:细粒度线程级并行性以实现单线程性能
批准号:
EP/W00576X/1
负责人:
Timothy Jones
金额:
$139.12万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --

项目摘要

项目成果

Timothy Jones的其他基金

相似基金

相关文献

中文摘要
翻译
自世纪之交以来,多核处理器在几乎所有计算领域都变得司空见惯。性能不再仅仅来自指令级并行性(ILP)的提取,它现在还要求软件开发人员或编译器将应用程序分解为多个指令流,以利用粗粒度线程级并行性(TLP)。虽然单线程性能对于一大类程序非常有利,但它仍然非常重要,特别是在应用程序的连续部分,其中执行速度可能决定整个程序的性能(有时被称为“Amdahl残酷定律”)。此外,单线程性能的提高使所有应用程序受益,因为每个线程都会经历性能提升,从而影响代码的所有部分-顺序和并行。然而,提高单线程性能是困难的。向多核的转变是由于提取ILP的复杂无序硬件方案的功率限制(由底层晶体管技术中的Dennard缩放失败引起)。虽然设计者仍然增加了无序的指令窗口,但不幸的是,这只会产生微小的差异,未来的设计预计将受到波拉克规则和ILP(ILP墙)的基本限制的限制。相反,尽管许多应用程序将从利用TLP中获得巨大的性能提升,但实际提取TLP仍然是一个挑战(John Hennessy说编写并行代码是“一个与计算机科学所面临的任何问题一样困难的问题”)。这个项目采取了一种截然不同的方法。它没有带着复杂的无序执行方案回到未来,而是探索了ILP和现代多核所利用的粗粒度TLP之间的空间。特别是,它专注于从核内和核之间的单个指令流中提取细粒度的TLP。一方面,它将研究对应用程序透明地识别和启动独立的短时间运行的线程(硬件线程)的方案,以提高单线程的性能。另一方面,它将研究编译器技术来表明这种并行性,硬件能够在多个紧密耦合的核心内和跨多个核心利用这种并行性。如果成功,该项目将在提高核心资源利用率和以可扩展方式增加这些资源的能力的推动下,导致高性能核心的性能发生阶段性变化。它还将开辟更广阔的设计空间,用无序的管道复杂性来换取ILP与增加的TLP,以在面积、效率和应用领域适宜性之间找到更好的平衡。
英文摘要
Since the turn of the century, multicore processors have become commonplace in almost all computing domains. Instead of performance coming solely from the extraction of instruction-level parallelism (ILP), it now also requires software developers or compilers to break applications into multiple streams of instructions to exploit coarse-grained thread-level parallelism (TLP). Whilst extremely beneficial for a large class of programs, single-threaded performance still matters greatly, especially during sequential parts of an application where execution speed can dominate overall program performance (sometimes dubbed "Amdahl's cruel law"). In addition, improvements in single-threaded performance benefit all applications, as each thread experiences a performance uplift, thus impacting all parts of the code-sequential and parallel.However, improving single-threaded performance is hard. The move to multicore was driven by the power limitations of complex out-of-order hardware schemes to extract ILP (caused by the failure of Dennard scaling in the underlying transistor technologies). While designers do still increase the out-of-order instruction window, unfortunately this only makes a marginal difference and future designs are expected to be limited by Pollack's rule and the fundamental limits of ILP (the ILP wall). Conversely, although many applications would see a major performance boost from taking advantage of TLP, actually extracting it remains a challenge (John Hennessy said writing parallel code is "a problem that's as hard as any that computer science has faced").This project takes a radically different approach. Instead of going back to the future with elaborate schemes for out-of-order execution, it explores the space between ILP and the coarse-grained TLP exploited by modern multicores. In particular, it focuses on the extraction of fine-grained TLP from a single stream of instructions within and across cores. On the one hand it will investigate schemes to identify and spin-up independent short-running threads (hardware threadlets) transparently to the application, so as to boost single-threaded performance. On the other, it will research compiler techniques to indicate this parallelism, with the hardware able to exploit it within and across multiple tightly coupled cores. If successful, this project would lead to a step change in performance of high-performance cores, driven by increased utilisation of core resources and the ability to increase those resources in a scalable manner. It would also open up a broader design space, trading out-of-order pipeline complexity for ILP with increased TLP, to find better balances between area, efficiency and application-domain suitability.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Decoupled Vector Runahead
解耦矢量超前运行
DOI: 10.1145/3613424.3614255
发表时间: 2023
期刊:
影响因子: --
作者: [Naithani A]
通讯作者: Naithani A
Vector Runahead for Indirect Memory Accesses
间接内存访问的向量超前运行
DOI: 10.1109/mm.2022.3163132
发表时间: 2022
期刊: IEEE Micro
影响因子: 3.6
作者: [Naithani A]
通讯作者: Naithani A
CAPcelerate: Capabilities for Heterogeneous Accelerators
  • 批准号:
    EP/V000381/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $153.2万
  • 财政年份:
    2020
  • 负责人:
    Timothy Jones
  • 依托单位:
Automatic Binary Parallelisation
  • 批准号:
    EP/P020011/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $108.33万
  • 财政年份:
    2017
  • 负责人:
    Timothy Jones
  • 依托单位:
Warwick MRC Proximity to Discovery - Industry Engagement Fund (WMIEF)
  • 批准号:
    MC_PC_15064
  • 项目类别:
    Intramural
  • 资助金额:
    $12.74万
  • 财政年份:
    2016
  • 负责人:
    Timothy Jones
  • 依托单位:
University of Warwick Experimental Equipment Proposal
  • 批准号:
    EP/M028186/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $65.72万
  • 财政年份:
    2015
  • 负责人:
    Timothy Jones
  • 依托单位:
海外基金