Managing Reconfigurable Instructions The end of Denard scaling has brought with it the end of performance improvements using traditional microprocess
Managing Reconfigurable Instructions The end of Denard scaling has brought with it the end of performance improvements using traditional microprocess
批准号:
2265128
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
未结题
起止时间:
2019 至 --
中文摘要
解决这一问题的一种方法是多核扩展,这种方法既不是没有挑战,也不会无限期地继续提供帮助。ASIC是另一种有希望的方法,它消除了传统微处理器设计的许多管理费用。然而,到目前为止,基于ASIC的加速器只存在于特定的域中,例如在计算分组检查序列的NIC中或在执行简单数学运算的FPGA中。这里的问题有三个:1。选择有效的加速器是一个挑战。一旦硬化,ASIC就不能改变-这就压缩了上面的困难。加速器不可能都靠近CPU,这意味着通信成本可能很高。可配置的硬件解决了这些问题。CGRA是保持ASIC的性能优势和FPGA的灵活性的一种方法。选择合适的硬块仍然是一个挑战,将现有程序映射到CGRA是困难的。硬件加速器的空间远远大于现有的CGRA,但需要更多的编译工作来支持各种硬件的CGRA。现有的商业解决方案,如Merlin编译器,主要集中在将计算卸载到FPGA上,并支持有限数量的固定CGRA体系结构。Legup也扮演着类似的角色,但目标是现场可编程门阵列。DSL也是编写CGRA的一种典型方式,但学习曲线很陡峭,通常适用于少数CGRA。现有的用于映射到具有硬块的设备的技术主要集中在固定功能设备上。我建议从遗留代码中以CGRA为目标,就像这些代码以固定函数加速器为目标一样。更广泛地说,我的工作将提供对可精细重构设计的广泛适用性和ASIC效率之间的权衡的理解。一个明显的问题是:目标是什么是最好的体系结构?解决这个问题的一种方法是:给定一些固定功能的ASIC,如何通过引入可自由配置来将加速器的覆盖范围扩展到许多应用?编译器在回答这个问题中起着关键作用,因为从多个应用程序卸载功能的问题变成了将不同应用程序映射到差别很小的加速器的任务。但这是一个鸡和蛋的问题:我们需要一种技术,在适当的编译器技术开发之前,从ASIC加速器转移到具有少量可重构性但适用性更广的加速器。但如果没有编译器技术,潜在的加速器覆盖好处将是未知的。我建议开发一个工具来解决这个问题。给出一个带有相关功能描述和一组应用的ASIC,它将找到一个具有适量可重构性的可重配置加速器。该项目有几个阶段:1.确定既能捕捉加速器行为,又能捕捉可重构性的合适机会的合适的功能描述。有几种语言可以用来描述CGRA,也应该作为起点使用,多伦多的CGRA-ME,图宾根的CGADL和U Washington的SPR。这一步在很大程度上需要考虑什么样的可重构性是有意义的,并理解一些更高级别的描述语言。基本上,我把这部分看作是“获得人类对我想要做的事情的直觉”。2.找到一种有效的方法来匹配这种灵活的加速器与代码库的代码,在这种方式下,我们将灵活性降至最低,同时最大化覆盖范围。这种方法的想法是输出具有“正确”灵活性级别的描述。3.最后一部分是将部分可重构的加速器与一些新代码进行匹配的PASS。
英文摘要
One approach to addressing this problemis multi-core scaling, an approach neither absent of challengesnor expected to continue to help indefinitely.ASICs, which remove many of the overheads of traditional microprocessordesign, are another promising approach. However, thus far, ASIC-basedaccelerators exist only in specific domains, for example in NICscomputing packet check sequences or in FPGAs performingsimple math operations. The problem here is threefold:1. Selection of effective accelerators is challenging.2. Once hardened, an ASIC cannot be changed --- this compacts the difficulties above.3. Accelerators cannot all be close to the CPU, meaning that communication costs can be high.Reconfigurable hardware addresses these concerns.CGRAs are one approach to maintaining the performance benefits of ASICs and the flexibility of FPGAs. Selection of suitable hard-blocks is still a challenge,and mapping existing programs to CGRAs is difficult. The space of hardware accelerators is far greater than existing CGRAs,but more compiler work is neededto support CGRAs with a variety of hardware. Existing commercialsolutions, such as the Merlin compiler focus largely onoffloading computation to FPGAs and support a limited number offixed CGRA architectures. LegUp fulfills a similarrole, but targeting FPGAs. DSLs are also a typical way to program CGRAs but comewith steep learning curves and typically apply to a small numberof CGRAs. Existing techniques for mapping to devices withhard-blocks focus largely on fixed functionalitydevices. I propose targeting CGRAsfrom legacy code in the same way that these works targetfixed-function accelerators. More broadly, my workwill provide understanding about the tradeoffs between the broadapplicability of finely reconfigurable designs and the efficiencyof ASICs.An obvious question is: what is the best architecture to target?One approach to this question is: given some fixed-function ASIC, how can thecoverage of the accelerator be extended to a number of applications by introducing somereconfigurability? The compilerplays a key role in answering this question, because the problem of offloading functionalityfrom multiple applications becomes a task of mapping different applications toaccelerators that differ very slightly.But this is a chicken-and-egg problem: we need a technique to move from anASIC accelerator to an accelerator with a small amount of reconfigurability but muchbroader applicability before the appropriate compiler techniques can bedeveloped. But without the compiler techniques, the potential acceleratorcoverage benefits will be unknown.I propose the development of a tool that will address this problem. Given anASIC with an associated functional description and a set of applications, it willfind a reconfigurable accelerator with an appropriate amount of reconfigurability.There are several phases to this project:1. Identify a suitable functional description that captures both accelerator behaviour, and suitable opportunities for reconfigurability. There are a few languages that can be used to describe CGRAs that should also be used as a starting point, Toronto's CGRA-ME, Tubingen's CGADL and U Washington's SPR. This step largely requires thinking about what kinds of reconfigurability make sense, and understanding a few higher level description languages. Basically, I see this part as the ``get a human intuition for what I want to do''. 2. Find an efficient way to match this flexible accelerator to code to a code-base in a way where we minimise flexibility while maximising coverage. The idea of this one is to output a description with the ``right'' level of flexibility. 3. The last part is a pass that can match the partially-reconfigurable accelerator to some new code.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金