Hardware-Aware Machine Learning: Modeling and Optimization

Hardware-Aware Machine Learning: Modeling and Optimization
复制标题

DOI:
10.1145/3240765.3243479
复制
发表时间:
2018-09
期刊:
2018 IEEE/ACM International Conference on Computer-Aided Design (ICCAD)
影响因子:
--
通讯作者:
Diana Marculescu;Dimitrios Stamoulis;E. Cai
Diana Marculescu;Dimitrios Stamoulis;E. Cai
中科院分区:
其他
文献类型:
--
作者:
Diana Marculescu;Dimitrios Stamoulis;E. Cai

文献摘要

相似文献

最近机器学习(ML)应用的突破,特别是深度学习(DL)的突破,使深度学习模型成为几乎每一个现代计算系统的关键组件。在各种平台(从移动设备到数据中心)上部署的DL应用程序越来越受欢迎,这导致了与硬件本身带来的限制相关的过多设计挑战。“深度神经网络(DNN)进行推理的延迟或能量成本是多少?”“有没有可能在对模型进行训练之前就预测这种延迟或能源消耗?”如果是,机器学习者如何利用这些模型来设计硬件最优的DNN进行部署?从延长移动设备的电池寿命到降低在云中执行的DL模型的运行时间要求,这些问题的答案引起了极大的关注。没有正确建模的东西是无法优化的。因此,甚至在训练模型之前,在用于进行推理的服务期间了解DL模型的硬件效率是重要的。这一关键观察结果促使人们使用预测模型来捕获ML应用程序的硬件性能或能效。此外,ML实践者目前面临着设计DNN模型的任务,即调整DNN体系结构的超参数,同时优化DL模型的准确性及其硬件效率。因此,最先进的方法提出了硬件感知的超参数优化技术。在本文中,我们对ML应用程序的硬件感知建模和优化方面的最新工作和精选结果进行了全面的评估。随着数字图书馆应用继续对相关硬件系统和平台产生重大影响,我们还重点介绍了几个悬而未决的问题,这些问题有望在未来几年产生新的硬件感知设计。
Recent breakthroughs in Machine Learning (ML) applications, and especially in Deep Learning (DL), have made DL models a key component in almost every modern computing system. The increased popularity of DL applications deployed on a wide-spectrum of platforms (from mobile devices to datacenters) have resulted in a plethora of design challenges related to the constraints introduced by the hardware itself. “What is the latency or energy cost for an inference made by a Deep Neural Network (DNN)?” “Is it possible to predict this latency or energy consumption before a model is even trained?” “If yes, how can machine learners take advantage of these models to design the hardware-optimal DNN for deployment?” From lengthening battery life of mobile devices to reducing the runtime requirements of DL models executing in the cloud, the answers to these questions have drawn significant attention. One cannot optimize what isn't properly modeled. Therefore, it is important to understand the hardware efficiency of DL models during serving for making an inference, before even training the model. This key observation has motivated the use of predictive models to capture the hardware performance or energy efficiency of ML applications. Furthermore, ML practitioners are currently challenged with the task of designing the DNN model, i.e., of tuning the hyper-parameters of the DNN architecture, while optimizing for both accuracy of the DL model and its hardware efficiency. Therefore, state-of-the-art methodologies have proposed hardware-aware hyper-parameter optimization techniques. In this paper, we provide a comprehensive assessment of state-of-the-art work and selected results on the hardware-aware modeling and optimization for ML applications. We also highlight several open questions that are poised to give rise to novel hardware-aware designs in the next few years, as DL applications continue to significantly impact associated hardware systems and platforms.