Managing Reconfigurable Instructions The end of Denard scaling has brought with it the end of performance improvements using traditional microprocess
Managing Reconfigurable Instructions The end of Denard scaling has brought with it the end of performance improvements using traditional microprocess
批准号:
2265128
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
未结题
起止时间:
2019 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
One approach to addressing this problemis multi-core scaling, an approach neither absent of challengesnor expected to continue to help indefinitely.ASICs, which remove many of the overheads of traditional microprocessordesign, are another promising approach. However, thus far, ASIC-basedaccelerators exist only in specific domains, for example in NICscomputing packet check sequences or in FPGAs performingsimple math operations. The problem here is threefold:1. Selection of effective accelerators is challenging.2. Once hardened, an ASIC cannot be changed --- this compacts the difficulties above.3. Accelerators cannot all be close to the CPU, meaning that communication costs can be high.Reconfigurable hardware addresses these concerns.CGRAs are one approach to maintaining the performance benefits of ASICs and the flexibility of FPGAs. Selection of suitable hard-blocks is still a challenge,and mapping existing programs to CGRAs is difficult. The space of hardware accelerators is far greater than existing CGRAs,but more compiler work is neededto support CGRAs with a variety of hardware. Existing commercialsolutions, such as the Merlin compiler focus largely onoffloading computation to FPGAs and support a limited number offixed CGRA architectures. LegUp fulfills a similarrole, but targeting FPGAs. DSLs are also a typical way to program CGRAs but comewith steep learning curves and typically apply to a small numberof CGRAs. Existing techniques for mapping to devices withhard-blocks focus largely on fixed functionalitydevices. I propose targeting CGRAsfrom legacy code in the same way that these works targetfixed-function accelerators. More broadly, my workwill provide understanding about the tradeoffs between the broadapplicability of finely reconfigurable designs and the efficiencyof ASICs.An obvious question is: what is the best architecture to target?One approach to this question is: given some fixed-function ASIC, how can thecoverage of the accelerator be extended to a number of applications by introducing somereconfigurability? The compilerplays a key role in answering this question, because the problem of offloading functionalityfrom multiple applications becomes a task of mapping different applications toaccelerators that differ very slightly.But this is a chicken-and-egg problem: we need a technique to move from anASIC accelerator to an accelerator with a small amount of reconfigurability but muchbroader applicability before the appropriate compiler techniques can bedeveloped. But without the compiler techniques, the potential acceleratorcoverage benefits will be unknown.I propose the development of a tool that will address this problem. Given anASIC with an associated functional description and a set of applications, it willfind a reconfigurable accelerator with an appropriate amount of reconfigurability.There are several phases to this project:1. Identify a suitable functional description that captures both accelerator behaviour, and suitable opportunities for reconfigurability. There are a few languages that can be used to describe CGRAs that should also be used as a starting point, Toronto's CGRA-ME, Tubingen's CGADL and U Washington's SPR. This step largely requires thinking about what kinds of reconfigurability make sense, and understanding a few higher level description languages. Basically, I see this part as the ``get a human intuition for what I want to do''. 2. Find an efficient way to match this flexible accelerator to code to a code-base in a way where we minimise flexibility while maximising coverage. The idea of this one is to output a description with the ``right'' level of flexibility. 3. The last part is a pass that can match the partially-reconfigurable accelerator to some new code.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金