Lightweight kernel isolation with virtualization and VM functions

Lightweight kernel isolation with virtualization and VM functions
复制标题

DOI:
10.1145/3381052.3381328
复制
发表时间:
2020-03
期刊:
Proceedings of the 16th ACM SIGPLAN/SIGOPS International Conference on Virtual Execution Environments
影响因子:
--
通讯作者:
Vikram Narayanan;Yongzhe Huang;Gang Tan;T. Jaeger;A. Burtsev
Vikram Narayanan;Yongzhe Huang;Gang Tan;T. Jaeger;A. Burtsev
中科院分区:
其他
文献类型:
--
作者:
Vikram Narayanan;Yongzhe Huang;Gang Tan;T. Jaeger;A. Burtsev

文献摘要

相似文献

商品操作系统在单个地址空间中执行核心内核子系统,同时还有数百个动态加载的扩展和设备驱动程序。内核中缺乏隔离意味着内核任何子系统或设备驱动程序中的漏洞都为对整个内核进行成功攻击开辟了道路。从历史上看,由于硬件隔离原语成本高昂,内核内的隔离一直难以实现。然而,最近的CPU带来了一系列新机制。带有虚拟机功能的扩展页表(EPT)切换和内存保护键(MPK)提供了内存隔离以及跨保护域边界的调用,其开销与系统调用相当。不幸的是,MPK和EPT切换都没有为特权环0内核代码的隔离提供架构支持,即对特权指令的控制以及在隔离域之间转换时安全恢复系统状态的明确入口点。我们的工作开发了一系列用于轻量级隔离特权内核代码的技术。为了控制特权指令的执行,我们依靠一个最小的虚拟机监控程序,它透明地将系统剥夺特权,使其成为非根VT - x客户机。我们开发了一个新的隔离边界,它利用带有VMFUNC指令的扩展页表(EPT)切换。我们定义了一组不变量,使我们能够在内核复杂的执行模型面前隔离内核组件,例如,提供可抢占的并发中断处理程序的隔离。为了最小化虚拟化的开销,我们开发了对跨隔离域无退出中断传递的支持。我们通过在Linux内核中开发几个设备驱动程序的隔离版本来评估我们的方法。
Commodity operating systems execute core kernel subsystems in a single address space along with hundreds of dynamically loaded extensions and device drivers. Lack of isolation within the kernel implies that a vulnerability in any of the kernel subsystems or device drivers opens a way to mount a successful attack on the entire kernel. Historically, isolation within the kernel remained prohibitive due to the high cost of hardware isolation primitives. Recent CPUs, however, bring a new set of mechanisms. Extended page-table (EPT) switching with VM functions and memory protection keys (MPKs) provide memory isolation and invocations across boundaries of protection domains with overheads comparable to system calls. Unfortunately, neither MPKs nor EPT switching provide architectural support for isolation of privileged ring 0 kernel code, i.e., control of privileged instructions and well-defined entry points to securely restore state of the system on transition between isolated domains. Our work develops a collection of techniques for lightweight isolation of privileged kernel code. To control execution of privileged instructions, we rely on a minimal hypervisor that transparently deprivileges the system into a non-root VT-x guest. We develop a new isolation boundary that leverages extended page table (EPT) switching with the VMFUNC instruction. We define a set of invariants that allows us to isolate kernel components in the face of an intricate execution model of the kernel, e.g., provide isolation of preemptable, concurrent interrupt handlers. To minimize overheads of virtualization, we develop support for exitless interrupt delivery across isolated domains. We evaluate our approach by developing isolated versions of several device drivers in the Linux kernel.