DNN-SAM: Split-and-Merge DNN Execution for Real-Time Object Detection

DNN-SAM: Split-and-Merge DNN Execution for Real-Time Object Detection
复制标题

DOI:
10.1109/rtas54340.2022.00021
复制
发表时间:
2022-05
期刊:
2022 IEEE 28th Real-Time and Embedded Technology and Applications Symposium (RTAS)
影响因子:
--
通讯作者:
Woo-Sung Kang;Siwoo Chung;Jeremy Yuhyun Kim;Youngmoon Lee;Kilho Lee;Jinkyu Lee;K. Shin;H. Chwa
Woo-Sung Kang;Siwoo Chung;Jeremy Yuhyun Kim;Youngmoon Lee;Kilho Lee;Jinkyu Lee;K. Shin;H. Chwa
中科院分区:
其他
文献类型:
--
作者:
Woo-Sung Kang;Siwoo Chung;Jeremy Yuhyun Kim;Youngmoon Lee;Kilho Lee;Jinkyu Lee;K. Shin;H. Chwa

文献摘要

被引文献

相似文献

由于自动汽车等实时对象检测系统需要处理从多个摄像头获取的输入图像,因此它们在提供通常基于机器学习(ML)的准确和及时的推断方面面临着重大挑战。为了应对这些挑战,我们希望为每个输入图像中具有不同关键性级别的不同部分提供不同级别的对象检测准确性和及时性。具体来说,我们开发了DNN-SAM,这是一个动态的分裂合并深度神经网络(DNN)执行和调度框架,可以为未修改的DNN模型实现无缝的分裂合并DNN执行。DNN-SAM不是在完整的DNN模型中处理一次整个输入图像,而是首先将DNN推理任务分成两个较小的子任务-一个强制性子任务专用于每个图像的安全关键(裁剪)部分,另一个可选的子任务用于处理缩小的图像-然后独立执行它们,最后将它们的结果合并为一个完整的推理。为了实现DNN-SAM对每个图像中对象的及时准确检测,我们还开发了两种调度算法,根据关键性级别对子任务进行优先级排序,并自适应地调整输入图像的规模以满足时间约束,同时最大限度地减少强制子任务的响应时间或最大限度地提高可选子任务的准确性。我们已经在一个有代表性的ML框架上实现并评估了DNN-SAM。我们的评估表明,DNN-SAM在安全关键区域的检测精度提高了2.0 -3.7\times$,平均推理延迟比现有方法降低了4.8 -9.7\times$,而不违反任何时间限制。
As real-time object detection systems, such as autonomous cars, need to process input images acquired from multiple cameras, they face significant challenges in delivering accurate and timely inferences often based on machine learning (ML). To meet these challenges, we want to provide different levels of object detection accuracy and timeliness to different portions within each input image with different criticality levels. Specifically, we develop DNN-SAM, a dynamic Split-And-Merge Deep Neural Network (DNN) execution and scheduling framework, that enables seamless split-and-merge DNN execution for unmodified DNN models. Instead of processing an entire input image once in a full DNN model, DNN-SAM first splits a DNN inference task into two smaller sub-tasks-a mandatory sub-task dedicated for a safety-critical (cropped) portion of each image and an optional sub-task for processing a down-scaled image–then executes them independently, and finally merges their results into a complete inference. To achieve DNN-SAM’s timely and accurate detection of objects in each image, we also develop two scheduling algorithms that prioritize sub-tasks according to their criticality levels and adaptively adjust the scale of the input image to meet the timing constraints while minimizing the response time of mandatory sub-tasks or maximizing the accuracy of optional sub-tasks. We have implemented and evaluated DNN-SAM on a representative ML framework. Our evaluation shows DNN-SAM to improve detection accuracy in the safety-critical region by $2.0-3.7\times$ and lower average inference latency by $4.8-9.7\times$ over existing approaches without violating any timing constraints.