big.VLITTLE: On-Demand Data-Parallel Acceleration for Mobile Systems on Chip

big.VLITTLE: On-Demand Data-Parallel Acceleration for Mobile Systems on Chip
复制标题

DOI:
10.1109/micro56248.2022.00025
复制
发表时间:
2022-10
期刊:
2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
T. Ta;Khalid Al-Hawaj;Nick Cebry;Yanghui Ou;Eric Hall;Courtney Golden;C. Batten
T. Ta;Khalid Al-Hawaj;Nick Cebry;Yanghui Ou;Eric Hall;Courtney Golden;C. Batten
中科院分区:
其他
文献类型:
--
作者:
T. Ta;Khalid Al-Hawaj;Nick Cebry;Yanghui Ou;Eric Hall;Courtney Golden;C. Batten

文献摘要

相似文献

单isa异构多核架构为在移动片上系统(soc)中执行任务并行工作负载提供了令人信服的高性能和高效率解决方案。除了任务并行工作负载之外,许多数据并行应用程序(如机器学习、计算机视觉和数据分析)也越来越多地在移动soc上运行,以提供实时用户交互。下一代可扩展的矢量架构,如RISC-V矢量扩展和Arm SVE,最近作为大型和小型系统的统一矢量抽象而出现。在本文中,我们提出了一种称为big的新型高效区域高性能架构。支持下一代矢量架构的VLITTLE,可有效加速传统大数据中的数据并行工作负载。小系统。大了。VLITTLE架构根据需要重新配置多个小内核,以便在执行数据并行工作负载时作为解耦的矢量引擎工作。我们的研究结果表明。VLITTLE系统可以在同等面积的大范围内实现1.6倍的性能加速。LITTLE系统配备了一个集成的矢量单元,跨多个数据并行应用程序,与任务并行工作负载的激进解耦矢量引擎相比,速度提高了1.7倍。
Single-ISA heterogeneous multi-core architectures offer a compelling high-performance and high-efficiency solution to executing task-parallel workloads in mobile systems on chip (SoCs). In addition to task-parallel workloads, many data-parallel applications, such as machine learning, computer vision, and data analytics, increasingly run on mobile SoCs to provide real-time user interactions. Next-generation scalable vector architectures, such as the RISC-V Vector Extension and Arm SVE, have recently emerged as unified vector abstractions for both large- and small-scale systems. In this paper, we propose novel area-efficient high-performance architectures called big.VLITTLE that support next-generation vector architectures to efficiently accelerate data-parallel workloads in conventional big.LITTLE systems. big.VLITTLE architectures reconFigure multiple little cores on demand to work as a decoupled vector engine when executing data-parallel workloads. Our results show that a big.VLITTLE system can achieve $1.6\times$ performance speedup over an area-comparable big.LITTLE system equipped with an integrated vector unit across multiple data-parallel applications and $1.7\times$ speedup compared to an aggressive decoupled vector engine for task-parallel workloads.