An Empirical Study of Common Challenges in Developing Deep Learning Applications

An Empirical Study of Common Challenges in Developing Deep Learning Applications
复制标题

DOI:
10.1109/issre.2019.00020
复制
发表时间:
2019-10
期刊:
2019 IEEE 30th International Symposium on Software Reliability Engineering (ISSRE)
影响因子:
--
通讯作者:
Tianyi Zhang;Cuiyun Gao;Lei Ma;Michael R. Lyu;Miryung Kim
Tianyi Zhang;Cuiyun Gao;Lei Ma;Michael R. Lyu;Miryung Kim
中科院分区:
其他
文献类型:
--
作者:
Tianyi Zhang;Cuiyun Gao;Lei Ma;Michael R. Lyu;Miryung Kim

文献摘要

被引文献

相似文献

深度学习的最新进展促进了许多智能系统和应用程序的创新,例如自动驾驶和图像识别。要寻求答案,本文在流行的问答网站堆栈溢出中介绍了深度学习问题的大规模经验研究。我们手动检查了715个问题的样本,并确定了七种常见问题。模型迁移和实施问题是仔细研究这些问题的最常见的三个问题,我们总结了五个主要根源可能受到研究社区的关注的原因,包括API MISSUSE,不正确的超参数选择,GPU计算,静态图表计算以及有限的调试和分析支持,我们的结果突出了对新技术的需求,例如改进软件差异测试。发展生产力和深度学习中的软件可靠性。
Recent advances in deep learning promote the innovation of many intelligent systems and applications such as autonomous driving and image recognition. Despite enormous efforts and investments in this field, a fundamental question remains under-investigated—what challenges do developers commonly face when building deep learning applications? To seek an answer, this paper presents a large-scale empirical study of deep learning questions in a popular Q&A website, Stack Overflow. We manually inspect a sample of 715 questions and identify seven kinds of frequently asked questions. We further build a classification model to quantify the distribution of different kinds of deep learning questions in the entire set of 39,628 deep learning questions. We find that program crashes, model migration, and implementation questions are the top three most frequently asked questions. After carefully examining accepted answers of these questions, we summarize five main root causes that may deserve attention from the research community, including API misuse, incorrect hyperparameter selection, GPU computation, static graph computation, and limited debugging and profiling support. Our results highlight the need for new techniques such as cross-framework differential testing to improve software development productivity and software reliability in deep learning.