On-demand-fork: a microsecond fork for memory-intensive and latency-sensitive applications

On-demand-fork: a microsecond fork for memory-intensive and latency-sensitive applications
复制标题

DOI:
10.1145/3447786.3456258
复制
发表时间:
2021-04
期刊:
Proceedings of the Sixteenth European Conference on Computer Systems
影响因子:
--
通讯作者:
Kaiyang Zhao;Sishuai Gong;Pedro Fonseca
Kaiyang Zhao;Sishuai Gong;Pedro Fonseca
中科院分区:
其他
文献类型:
--
作者:
Kaiyang Zhao;Sishuai Gong;Pedro Fonseca

文献摘要

被引文献

相似文献

Fork一直是Unix的进程创建系统调用。一开始,fork被誉为一个高效的系统调用,因为它在父进程和子进程之间共享的内存上使用了写时复制。然而,应用程序内存需求自早期以来急剧增加,并且fork简单地设置虚拟内存(例如,复制页表)现在是一个问题,即使对于只需要数百MB内存的应用程序也是如此。在实践中,分叉性能已经在一系列分叉大型进程的用例中阻碍了系统效率和延迟,例如容错系统,无服务器框架和测试框架。本文提出了按需fork,fork系统调用的快速实现,专门为具有大内存占用的应用程序而设计。按需分叉依赖于这样一种观察,即写时复制可以推广到页表,即使是在商品硬件上。按需fork比传统的fork执行速度更快,因为它在fork时在父代和子代之间额外共享页表,并在处理页面错误时按需选择性地复制小块页表。On-demand-fork是fork的直接替代品,不需要对应用程序或硬件进行更改。我们在一系列微基准测试和真实工作负载上评估了按需分叉。按需fork显著减少了fork调用时间,并提高了可扩展性。对于具有1 GB分配内存的进程,按需fork比Fork具有65倍的性能优势。我们还对知名应用程序的测试,模糊和快照工作负载进行了按需分叉评估,获得了59%至226%的执行吞吐量改进和高达99%的调用延迟减少。
Fork has long been the process creation system call for Unix. At its inception, fork was hailed as an efficient system call due to its use of copy-on-write on memory shared between parent and child processes. However, application memory demand has increased drastically since the early days and the cost incurred by fork to simply set up virtual memory (e.g., copy page tables) is now a concern, even for applications that only require hundreds of MBs of memory. In practice, fork performance already holds back system efficiency and latency across a range of uses cases that fork large processes, such as fault-tolerant systems, serverless frameworks, and testing frameworks. This paper proposes On-demand-fork, a fast implementation of the fork system call specifically designed for applications with large memory footprints. On-demand-fork relies on the observation that copy-on-write can be generalized to page tables, even on commodity hardware. On-demand-fork executes faster than the traditional fork implementation by additionally sharing page tables between parent and child at fork time and selectively copying page tables in small chunks, on-demand, when handling page faults. On-demand-fork is a drop-in replacement for fork that requires no changes to applications or hardware. We evaluated On-demand-fork on a range of micro-benchmarks and real-world workloads. On-demand-fork significantly reduces the fork invocation time and has improved scalability. For processes with 1 GB of allocated memory, On-demand-fork has a 65× performance advantage over Fork. We also evaluated On-demand-fork on testing, fuzzing, and snapshotting workloads of well-known applications, obtaining execution throughput improvements between 59% and 226% and up to 99% invocation latency reduction.