Serverless Data Analytics with Flint

Serverless Data Analytics with Flint
复制标题

DOI:
10.1109/cloud.2018.00063
复制
发表时间:
2018-03
期刊:
2018 IEEE 11th International Conference on Cloud Computing (CLOUD)
影响因子:
--
通讯作者:
Youngbin Kim;Jimmy J. Lin
Youngbin Kim;Jimmy J. Lin
中科院分区:
其他
文献类型:
--
作者:
Youngbin Kim;Jimmy J. Lin

文献摘要

被引文献

相似文献

围绕松耦合函数调用组织的无服务器架构代表了许多应用程序的新兴设计。最近的工作主要集中在面向用户的产品和事件驱动的处理管道。在本文中,我们探索了应用程序空间的一个完全不同的部分,并研究了使用无服务器架构对大数据进行分析处理的可行性。我们介绍Flint,这是一个原型Spark执行引擎,它利用AWS Lambda提供纯按需付费的成本模型。使用Flint,开发人员可以像以前一样使用PySpark,但不需要实际的Spark集群。我们描述了Flint的设计、实现和性能,沿着与无服务器分析相关的挑战。
Serverless architectures organized around loosely-coupled function invocations represent an emerging design for many applications. Recent work mostly focuses on user-facing products and event-driven processing pipelines. In this paper, we explore a completely different part of the application space and examine the feasibility of analytical processing on big data using a serverless architecture. We present Flint, a prototype Spark execution engine that takes advantage of AWS Lambda to provide a pure pay-as-you-go cost model. With Flint, a developer uses PySpark exactly as before, but without needing an actual Spark cluster. We describe the design, implementation, and performance of Flint, along with the challenges associated with serverless analytics.