An Empirical Study Towards Characterizing Deep Learning Development and Deployment Across Different Frameworks and Platforms

An Empirical Study Towards Characterizing Deep Learning Development and Deployment Across Different Frameworks and Platforms
复制标题

DOI:
10.1109/ase.2019.00080
复制
发表时间:
2019-09
期刊:
2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE)
影响因子:
--
通讯作者:
Qianyu Guo;Sen Chen;Xiaofei Xie;Lei Ma;Q. Hu;Hongtao Liu;Yang Liu;Jianjun Zhao;Xiaohong Li
Qianyu Guo;Sen Chen;Xiaofei Xie;Lei Ma;Q. Hu;Hongtao Liu;Yang Liu;Jianjun Zhao;Xiaohong Li
中科院分区:
其他
文献类型:
--
作者:
Qianyu Guo;Sen Chen;Xiaofei Xie;Lei Ma;Q. Hu;Hongtao Liu;Yang Liu;Jianjun Zhao;Xiaohong Li

文献摘要

被引文献

相似文献

深度学习(DL)最近取得了巨大的成功。各种DL框架和平台在促进此类进步方面起着关键作用。但是,现有框架和平台的体系结构设计和实现的差异为DL软件开发和部署带来了新的挑战。到目前为止,还没有关于各种主流框架和平台如何影响DL软件开发和实践部署的研究。为了填补这一空白,我们迈出了第一步,朝着了解最广泛使用的DL框架和平台如何支持DL软件开发和部署。我们通过使用两种类型的DNN体系结构和三个流行数据集对这些框架和平台进行了系统的研究。 (1)对于开发过程,我们研究了在相同的运行时训练配置或相同模型权重/偏见下的预测准确性。我们还通过利用现有的对抗攻击技术来研究训练有素的模型的对抗性鲁棒性。实验结果表明,跨框架的计算差异可能导致明显的预测准确性下降,这应该引起DL开发人员的注意。 (2)对于部署过程,我们研究了预测准确性和性能(指时成本和内存消耗)当训练有素的模型从PC迁移/量化为实际的移动设备和Web浏览器时。 DL平台研究揭示了迁移和量化仍然存在兼容性和可靠性问题。同时,我们通过使用结果作为基准来找到几个DL软件错误。我们通过利益相关者和工业积极反馈的错误确认进一步验证结果,以突出我们的研究含义。通过我们的研究,我们总结了实用准则,确定挑战并查明新的研究方向,例如了解DL框架和平台的特征,避免兼容性和可靠性问题,检测DL软件错误以及减少时间成本和记忆消耗,以开发和部署高质量的DL系统有效。
Deep Learning (DL) has recently achieved tremendous success. A variety of DL frameworks and platforms play a key role to catalyze such progress. However, the differences in architecture designs and implementations of existing frameworks and platforms bring new challenges for DL software development and deployment. Till now, there is no study on how various mainstream frameworks and platforms influence both DL software development and deployment in practice. To fill this gap, we take the first step towards understanding how the most widely-used DL frameworks and platforms support the DL software development and deployment. We conduct a systematic study on these frameworks and platforms by using two types of DNN architectures and three popular datasets. (1) For development process, we investigate the prediction accuracy under the same runtime training configuration or same model weights/biases. We also study the adversarial robustness of trained models by leveraging the existing adversarial attack techniques. The experimental results show that the computing differences across frameworks could result in an obvious prediction accuracy decline, which should draw the attention of DL developers. (2) For deployment process, we investigate the prediction accuracy and performance (refers to time cost and memory consumption) when the trained models are migrated/quantized from PC to real mobile devices and web browsers. The DL platform study unveils that the migration and quantization still suffer from compatibility and reliability issues. Meanwhile, we find several DL software bugs by using the results as a benchmark. We further validate the results through bug confirmation from stakeholders and industrial positive feedback to highlight the implications of our study. Through our study, we summarize practical guidelines, identify challenges and pinpoint new research directions, such as understanding the characteristics of DL frameworks and platforms, avoiding compatibility and reliability issues, detecting DL software bugs, and reducing time cost and memory consumption towards developing and deploying high quality DL systems effectively.