Fast inference services for alternative deep learning structures

Fast inference services for alternative deep learning structures
复制标题

DOI:
10.1145/3318216.3363331
复制
发表时间:
2019-11
期刊:
Proceedings of the 4th ACM/IEEE Symposium on Edge Computing
影响因子:
--
通讯作者:
Eduardo Romero;Christopher Stewart;Nathaniel Morris
Eduardo Romero;Christopher Stewart;Nathaniel Morris
中科院分区:
其他
文献类型:
--
作者:
Eduardo Romero;Christopher Stewart;Nathaniel Morris

文献摘要

相似文献

人工智能推理服务接收请求、对数据进行分类并快速响应。这些服务是人工智能驱动的物联网、推荐引擎和视频分析的基础。神经网络被广泛使用,因为它们提供准确的结果和快速的推理,但很难解释它们的分类。基于树的深度学习模型可以提供准确性,并且天生可解释。然而,由于分支误预测和缓存未命中会导致执行效率低下,因此很难实现高推理率。我的研究旨在基于树模型产生低延迟的推理服务。我将利用大型L3缓存的出现,将基于树的模型推理从顺序分支转换为快速的缓存内查找。我们的方法从经过充分训练的精确的基于树的模型开始,编译它们以在目标处理器上进行推理,并有效地执行推理。如果成功,我们的方法将使人工智能服务取得质的进步。基于树的模型可以在一次通过中报告分类中最重要的特征。相比之下,神经网络需要迭代方法来解释其结果。考虑交互式人工智能推荐服务,其中用户寻求明确排序他们的即时偏好以吸引首选内容。基于树的模型可以比神经网络更快地提供用户反馈。基于树的模型也比神经网络具有更少的预测方差。给定相同的训练数据,神经网络需要许多推理来量化边界分类的方差。基于树的快速推理可以解释秒(相对于分钟)的差异。我们的方法表明,竞争机器学习方法可以提供相当的准确性,但需要完全不同的架构和平台支持。
AI inference services receive requests, classify data and respond quickly. These services underlie AI-driven Internet of Things, recommendation engines and video analytics. Neural networks are widely used because they provide accurate results and fast inference, but it is hard to explain their classifications. Tree-based deep learning models can provide accuracy and are innately explainable. However, it is hard to achieve high inference rates because branch misprediction and cache misses produce inefficient executions. My research seeks to produce low latency inference services based on tree-based models. I will exploit the emergence of large L3 caches to convert tree-based model inference from sequential branching toward fast, in-cache lookups. Our approach begins with fully trained, accurate tree-based models, compiles them for inference on target processors and executes inference efficiently. If successful, our approach will enable qualitative advances in AI services. Tree-based models can report the most significant features in a classification in a single pass. In contrast, neural networks require iterative approaches to explain their results. Consider interactive AI recommendation services where users seek to explicitly order their instantaneous preferences to attract preferred content. Tree-based models can provide user feedback much more quickly than neural networks. Tree-based models also have less prediction variance than neural networks. Given the same training data, neural networks require many inferences to quantify variances of borderline classifications. Fast tree-based inference can explain variance in seconds (versus minutes). Our approach shows that competing machine learning approaches can provide comparable accuracy but desire wholly different architectural and platform support.