Integrating algorithmic parameters into benchmarking and design space exploration in 3D scene understanding

Integrating algorithmic parameters into benchmarking and design space exploration in 3D scene understanding
复制标题

DOI:
10.1145/2967938.2967963
复制
发表时间:
2016-09
期刊:
2016 International Conference on Parallel Architecture and Compilation Techniques (PACT)
影响因子:
--
通讯作者:
Sreekar Shenoy;Bruno Bodin;Luigi Nardi;M. Zia;Harry Wagstaff;Govind Sreekar Shenoy;M. Emani;John Mawer;Christos Kotselidis;A. Nisbet;M. Luján;Björn Franke;P. Kelly;M. O’Boyle
Sreekar Shenoy;Bruno Bodin;Luigi Nardi;M. Zia;Harry Wagstaff;Govind Sreekar Shenoy;M. Emani;John Mawer;Christos Kotselidis;A. Nisbet;M. Luján;Björn Franke;P. Kelly;M. O’Boyle
中科院分区:
其他
文献类型:
--
作者:
Sreekar Shenoy;Bruno Bodin;Luigi Nardi;M. Zia;Harry Wagstaff;Govind Sreekar Shenoy;M. Emani;John Mawer;Christos Kotselidis;A. Nisbet;M. Luján;Björn Franke;P. Kelly;M. O’Boyle

文献摘要

相似文献

系统设计人员通常使用经过充分研究的基准来评估和改进新的架构和编译器。我们基于昨天的应用设计明天的系统。在本文中,我们调查一个新兴的应用程序,三维场景的理解,可能是显着的移动的空间在不久的将来。到目前为止,这个应用程序只能在桌面GPU上实时运行。在这项工作中,我们研究如何将其映射到功率受限的嵌入式系统。我们的方法的关键是增量协同设计探索的想法,其中关注域层的优化选择与低级编译器和架构选择一起增量探索。这种探索的目标是减少执行时间,同时最大限度地减少功耗并满足我们的结果质量目标。由于设计空间太大,无法彻底评估,我们使用基于随机森林预测器的主动学习来找到好的设计。我们表明,我们的方法可以,第一次,实现密集的3D映射和跟踪的实时范围内的一个流行的嵌入式设备上的1W的功率预算。与现有技术相比,执行时间缩短了4.8倍,功耗降低了2.8倍。
System designers typically use well-studied benchmarks to evaluate and improve new architectures and compilers. We design tomorrow's systems based on yesterday's applications. In this paper we investigate an emerging application, 3D scene understanding, likely to be significant in the mobile space in the near future. Until now, this application could only run in real-time on desktop GPUs. In this work, we examine how it can be mapped to power constrained embedded systems. Key to our approach is the idea of incremental co-design exploration, where optimization choices that concern the domain layer are incrementally explored together with low-level compiler and architecture choices. The goal of this exploration is to reduce execution time while minimizing power and meeting our quality of result objective. As the design space is too large to exhaustively evaluate, we use active learning based on a random forest predictor to find good designs. We show that our approach can, for the first time, achieve dense 3D mapping and tracking in the real-time range within a 1W power budget on a popular embedded device. This is a 4.8× execution time improvement and a 2.8× power reduction compared to the state-of-the-art.