λ-NIC: Interactive Serverless Compute on SmartNICs
λ-NIC: Interactive Serverless Compute on SmartNICs
复制标题
λ-NIC:SmartNIC 上的交互式无服务器计算
DOI:
10.1145/3342280.3342341
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
M. Rosenblum
中科院分区:
文献类型:
--
作者:
Sean Choi;M. Shahbaz;B. Prabhakar;M. Rosenblum
1 OVERVIEW Serverless compute is emerging as an attractive cloud computing model that lets developers focus only on the core applications, building them as small, fine-grained workloads (i.e., lambdas), without having to worry about building and/or managing the infrastructure they run on. Cloud providers dynamically provision, deploy, patch, and monitor the infrastructure and its resources (e.g., compute, storage, memory, and network) for these workloads; with tenants only paying for the resources they consume at millisecond increments. The cloud providers generally put a strict limit on the compute time and resource that can be consumed by a single workload, in order to ensure that they can easily deploy and scale each workload without impacting the availability of other workloads. Thus, the workloads are short-lived with strict compute time and memory limits (up to 15 minutes and 3 GB, respectively, for Amazon Lambda [6]) and are often latency sensitive. Some examples of these workloads include real-time stream processing and generic API endpoints. Today, all major cloud vendors offer some form of serverless frameworks (Figure 1), such as Amazon Lambda [3], Google Cloud Functions [9], and Microsoft Azure Functions [7], along with opensource developments like OpenFaaS [13] and OpenWhisk [2]. These frameworks rely on virtualization and containers [10] to execute and scale tenants’ lambdas. These technologies were designed to maximize utilization of the providers’ physical infrastructure, while presenting each tenant with its own view of a completely isolated machine. With serverless computing, where server management is hidden from tenants, these virtualization technologies become redundant, unnecessarily bloating the code size of serverless workloads, and causing processing delays (of hundreds of milliseconds)