Sommelier: Curating DNN Models for the Masses

Sommelier: Curating DNN Models for the Masses
复制标题

DOI:
10.1145/3514221.3526173
复制
发表时间:
2022-06
期刊:
Proceedings of the 2022 International Conference on Management of Data
影响因子:
--
通讯作者:
Peizhen Guo;Bo Hu;Wenjun Hu
Peizhen Guo;Bo Hu;Wenjun Hu
中科院分区:
其他
文献类型:
--
作者:
Peizhen Guo;Bo Hu;Wenjun Hu

文献摘要

相似文献

深度学习模型存储库在当今的机器学习生态系统中是不可或缺的,以促进模型重用。然而,现有的模型存储库提供了一个用于模型检索的基本接口。用户有责任从潜在的数百种选择中进行分析和选择,这几乎不会减轻普通用户首先设计模型所需的专业知识。在本文中,我们介绍了Sommelier,这是一个位于典型DNN模型存储库之上的索引和查询系统,可以直接与推理服务或其他用例进行交互。给定推理任务类别的理想准确性目标和资源预算,Sommelier自动在存储库中搜索最合适的模型,而不需要用户手动分析。受手动迭代模型搜索过程和生成模型变体或具有公共段的模型的典型模型设计策略的启发,Sommelier基于其语义相关性(定义为模型产生相同结果的概率)组织DNN模型。这进一步与基于相对资源消耗的资源索引相结合。Sommelier作为一个独立的查询引擎实现,可以与现有的存储库(如TF-Hub)进行交互。TF-Hub中的163个模型的案例研究突出了不同模型系列之间的模型相关性程度,表明最佳候选模型可以轻松逃避手动分析。广泛的评估表明,Sommelier返回超过95%的查询的理想模型;当与推理服务器连接时,Sommelier可以通过自动模型切换将推理任务的第90百分位数尾部延迟降低6倍,远远超过典型的横向扩展系统优化。
Deep learning model repositories are indispensable in machine learning ecosystems today to facilitate model reuse. However, existing model repositories provide a bare-bone interface for model retrieval. The onus is on the user to profile and select from potentially hundreds of choices, barely relieving an average user of the expertise required to design the model in the first place. In this paper, we present Sommelier, an indexing and query system above typical DNN model repositories to interface directly with inference serving or other use cases. Given a desirable accuracy target and resource budget for an inference task category, Sommelier automatically searches through the repository for the most suitable model, without requiring manual profiling from the user. Motivated by manual iterative model search processes and typical model design strategies that generate model variants or models with common segments, Sommelier organizes DNN models based on their semantic correlation, defined as the probability of models producing the same results. This is further combined with a resource index based on relative resource consumption. Sommelier is implemented as a standalone query engine that can interface with an existing repository such as TF-Hub. A case study of 163 models in TF-Hub highlights the extent of model correlation across different model series, suggesting the best candidate model can easily evade manual profiling. Extensive evaluation shows that Sommelier returns the ideal model for over 95% of the queries; When interfaced with an inference server, Sommelier can reduce the 90th percentile tail latency of inference tasks by a factor of 6 via automatic model switching, far more than typical scale-out system optimizations.