DLHub: Simplifying publication, discovery, and use of machine learning models in science
DLHub: Simplifying publication, discovery, and use of machine learning models in science
复制标题
DOI:
10.1016/j.jpdc.2020.08.006
复制
发表时间:
2021-01-01
影响因子:
3.8
通讯作者:
Foster, Ian
中科院分区:
文献类型:
--
作者:
Li, Zhuozhao;Chard, Ryan;Foster, Ian
Machine Learning (ML) has become a critical tool enabling new methods of analysis and driving deeper understanding of phenomena across scientific disciplines. There is a growing need for "learning systems" to support various phases in the ML lifecycle. While others have focused on supporting model development, training, and inference, few have focused on the unique challenges inherent in science, such as the need to publish and share models and to serve them on a range of available computing resources. In this paper, we present the Data and Learning Hub for science (DLHub), a learning system designed to support these use cases. Specifically, DLHub enables publication of models, with descriptive metadata, persistent identifiers, and flexible access control. It packages arbitrary models into portable servable containers, and enables low-latency, distributed serving of these models on heterogeneous compute resources. We show that DLHub supports low-latency model inference comparable to other model serving systems including TensorFlow Serving, SageMaker, and Clipper, and improved performance, by up to 95%, with batching and memoization enabled. We also show that DLHub can scale to concurrently serve models on 500 containers. Finally, we describe five case studies that highlight the use of DLHub for scientific applications. (C) 2020 Elsevier Inc. All rights reserved.