Towards a Machine Learning-Assisted Kernel with LAKE

Towards a Machine Learning-Assisted Kernel with LAKE
复制标题

DOI:
10.1145/3575693.3575697
复制
发表时间:
2023-01
期刊:
Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2
影响因子:
--
通讯作者:
Henrique Fingler;Isha Tarte;Hangchen Yu;Ariel Szekely;Bodun Hu;Aditya Akella;C. Rossbach
Henrique Fingler;Isha Tarte;Hangchen Yu;Ariel Szekely;Bodun Hu;Aditya Akella;C. Rossbach
中科院分区:
其他
文献类型:
--
作者:
Henrique Fingler;Isha Tarte;Hangchen Yu;Ariel Szekely;Bodun Hu;Aditya Akella;C. Rossbach

文献摘要

相似文献

现代操作系统(OS)的复杂性,硬件的快速多样化以及机器学习(ML)的稳步发展促使我们探索ML改善OS内核决策的潜力。我们推测ML可以更好地管理子系统的权衡空间,例如内存管理和进程以及I/O调度,这些子系统目前依赖于手动调整的算法来提供合理的平均性能。我们探讨了在五个内核子系统中用ML驱动的决策来替换逻辑学,考虑了内核设计,共享操作系统级组件和硬件加速的影响。我们确定了障碍,解决了挑战,并描述了ML可以在内核空间中提供的好处的权衡。我们发现,使用GPU等专用硬件对于吸收ML决策所需的额外计算负载至关重要,但内核空间中加速器的可访问性较差是采用的障碍。我们还发现,ML和加速对OS的好处取决于子系统、工作负载和硬件,这表明在内核中使用ML将需要框架来帮助内核开发人员导航新的权衡空间。我们通过构建一个名为LAKE的系统来支持ML并在内核空间中暴露加速器来解决这些挑战。LAKE包括用于跨抽象层和模块边界的特性收集和管理的API。LAKE提供了用于管理加速的可变收益的机制,以及用于减轻用户和内核空间之间的资源争用的接口。我们表明,ML支持的I/O延迟预测器可以将其推理时间缩短高达96%。
The complexity of modern operating systems (OSes), rapid diversification of hardware, and steady evolution of machine learning (ML) motivate us to explore the potential of ML to improve decision-making in OS kernels. We conjecture that ML can better manage tradeoff spaces for subsystems such as memory management and process and I/O scheduling that currently rely on hand-tuned heuristics to provide reasonable average-case performance. We explore the replacement of heuristics with ML-driven decision-making in five kernel subsystems, consider the implications for kernel design, shared OS-level components, and access to hardware acceleration. We identify obstacles, address challenges and characterize tradeoffs for the benefits ML can provide that arise in kernel-space. We find that use of specialized hardware such as GPUs is critical to absorbing the additional computational load required by ML decisioning, but that poor accessibility of accelerators in kernel space is a barrier to adoption. We also find that the benefits of ML and acceleration for OSes is subsystem-, workload- and hardware-dependent, suggesting that using ML in kernels will require frameworks to help kernel developers navigate new tradeoff spaces. We address these challenge by building a system called LAKE for supporting ML and exposing accelerators in kernel space. LAKE includes APIs for feature collection and management across abstraction layers and module boundaries. LAKE provides mechanisms for managing the variable profitability of acceleration, and interfaces for mitigating contention for resources between user and kernel space. We show that an ML-backed I/O latency predictor can have its inference time reduced by up to 96% with acceleration.