Vista: Optimized System for Declarative Feature Transfer from Deep CNNs at Scale

Vista: Optimized System for Declarative Feature Transfer from Deep CNNs at Scale
复制标题

DOI:
10.1145/3318464.3389709
复制
发表时间:
2020-05
期刊:
Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data
影响因子:
--
通讯作者:
Supun Nakandala;Arun Kumar
Supun Nakandala;Arun Kumar
中科院分区:
其他
文献类型:
--
作者:
Supun Nakandala;Arun Kumar

文献摘要

被引文献

相似文献

用于机器学习的可扩展系统(ML)在很大程度上被孤立在数据流系统中,以用于无组织数据的结构化数据和深度学习系统。这个差距已经留下了工作量,可以共同分析两种形式的数据,并提供较差的系统支持,从而导致系统效率低下和用户的咕unt工作。我们弥合了一系列重要的工作负载类别:从深卷积神经网络(CNN)的特征转移,用于分析图像以及结构化数据。如今,在可扩展数据流和深度学习系统上执行功能传输面临两个关键系统问题:由于记忆不善而导致的冗余计算和崩溃 - 主持性效率低下。我们提出了Vista,这是一种新的数据系统,可以通过将此工作量提升到数据流和深度学习系统的声明级别来解决这些问题。 Vista会自动优化此工作负载的配置和执行,以减少计算冗余和工作负载崩溃的可能性。实际数据集上的实验表明,除了使功能传输更容易,Vista避免了工作量崩溃,并且与基准相比,Vista避免了工作负载崩溃,并将运行时间降低了58%至92%。
Scalable systems for machine learning (ML) are largely siloed into dataflow systems for structured data and deep learning systems for unstructured data. This gap has left workloads that jointly analyze both forms of data with poor systems support, leading to both low system efficiency and grunt work for users. We bridge this gap for an important class of such workloads: feature transfer from deep convolutional neural networks (CNNs) for analyzing images along with structured data. Executing feature transfer on scalable dataflow and deep learning systems today faces two key systems issues: inefficiency due to redundant computations and crash-proneness due to mismanaged memory. We present Vista, a new data system that resolves these issues by elevating this workload to a declarative level on top of dataflow and deep learning systems. Vista automatically optimizes the configuration and execution of this workload to reduce both computational redundancy and the potential for workload crashes. Experiments on real datasets show that apart from making feature transfer easier, Vista avoids workload crashes and reduces runtimes by 58% to 92% compared to baselines.