Skip to main content
  1. Index/

KVM/QEMU

Table of Contents

KVM (Kernel-based Virtual Machine) and QEMU (Quick Emulator) solve different halves of the same problem and are almost always used together. KVM is a Linux kernel module that turns the kernel into a hypervisor: with Intel VT-x or AMD-V, guest CPUs run on real hardware at near-native speed and guest memory is managed through extended page tables. KVM has no device model of its own — no disk, network, or firmware emulation. QEMU supplies that in userspace (virtio, USB, VGA, ACPI) and controls KVM by issuing ioctl calls on /dev/kvm — creating VMs, mapping memory, running vCPUs via KVM_RUN — while QMP exposes external management over a Unix socket. libvirt, originally from Red Hat, sits above both: it translates domain XML into QEMU command lines and lifecycle operations. The stack is the default on Linux: RHEL ships qemu-kvm, OpenShift Virtualization runs KubeVirt (virt-launcher → libvirt → QEMU), and OpenShift Sandboxed Containers uses Kata Containers with the same hypervisor to isolate pods in micro-VMs.

The main security concern is not KVM’s kernel module — relatively small — but QEMU’s device emulation: large C codebases parsing untrusted guest input, historically a source of VM escapes (e.g. VENOM, CVE-2015-3456). Defence in depth matters: run QEMU as the unprivileged qemu user, prefer virtio over legacy devices, and layer SELinux sVirt (per-VM MCS labels on QEMU and disk images via libvirt), seccomp (syscall and ioctl allowlists on /dev/kvm, vhost, VFIO), and cgroups. KubeVirt adds pod-level isolation; Kata adds a VM boundary so a container escape must cross a guest kernel first. Confidential VMs (SEV-SNP, TDX) go further: guest memory is encrypted so a compromised hypervisor cannot read pod or VM contents in plaintext — the path used by CoCo and Peer Pods on OpenShift.

Red Hat’s supported path remains KVM/QEMU via libvirt, not minimal VMMs like Firecracker (AWS) or Cloud Hypervisor (Kata upstream options with smaller device models). That choice trades a larger emulation surface for the breadth KubeVirt needs (TPM, GPU passthrough, live migration) and what CoCo needs (virtio-fs, vTPM, TEE launch ioctls). Operational hygiene: keep qemu-kvm and libvirt current via RHEL errata, audit domain XML for unnecessary legacy devices, enforce SELinux on RHCOS nodes, and use Trustee attestation where the threat model includes a hostile infrastructure operator. See also TCB, libvirt, Kata Containers, and Confidential VM.

Additional Information
#

Related

libvirt

libvirt is an open-source library, daemon, and toolset that provides a unified, stable API for managing virtualisation infrastructure — virtual machines, storage volumes, virtual networks, and host devices — across multiple hypervisor backends. It was originally written by Daniel Berrange at Red Hat and has since become the standard virtualisation management layer on Linux, underpinning KubeVirt, OpenStack Nova, oVirt/RHEV, Proxmox, and the virsh / virt-manager administrative tools. The core value proposition is hypervisor abstraction: the same libvirt API call creates a VM on KVM/QEMU, Xen, LXC, or (historically) VMware ESXi, without the management layer caring about the underlying implementation. In practice, the KVM/QEMU driver is the dominant use case on Linux; the others are progressively less-maintained but remain supported. libvirt communicates with hypervisors through driver-specific mechanisms — with QEMU, it generates the full QEMU command line from the domain XML definition and manages the QEMU process lifecycle, communicating with the running VM through the QEMU Monitor Protocol (QMP) over a Unix socket.

RHCOS (Red Hat Enterprise Linux CoreOS)

RHCOS (Red Hat Enterprise Linux CoreOS) is the operating system that runs on every OpenShift control plane and worker node. It is not a general-purpose Linux distribution — it is a purpose-built, immutable, container-optimised OS designed to run exclusively as a managed node in an OpenShift cluster. Its security posture is architecturally different from a hardened RHEL installation: rather than hardening a mutable system through configuration management, RHCOS makes the OS layer structurally resistant to modification by design. The root filesystem’s /usr tree is read-only (enforced at mount time by rpm-ostree and, in recent versions, by composefs over the OSTree object store), /etc and /var are writable but managed exclusively by the Machine Config Operator (MCO), and no package manager is available at runtime for ad-hoc software installation. An operator who wants to change any node-level configuration — kernel arguments, sysctl settings, systemd units, certificates, kubelet configuration — creates a MachineConfig object in the OpenShift API; the MCO renders it into an Ignition config, applies it to the target MachineConfigPool (master, worker, or custom), and drains and reboots the affected nodes in a rolling fashion. Direct SSH access to nodes for configuration changes is explicitly unsupported and actively discouraged — oc debug node/<name> is the supported emergency access path, dropping into a privileged container on the node’s host namespaces under audit.

KubeVirt

KubeVirt is a CNCF project that makes Kubernetes a native hypervisor management plane, allowing KVM virtual machines to be declared, scheduled, and operated through the Kubernetes API without a separate virtualisation management layer. The motivating use case is organisational convergence: teams running a mix of legacy VM workloads and modern containerised services no longer need two separate platforms (an OpenStack or vSphere cluster for VMs, a Kubernetes cluster for containers) with separate networking, storage, RBAC, and CI/CD integration. With KubeVirt, both workload types live in the same cluster, share the same kubectl and GitOps tooling, and are subject to the same scheduling, resource quota, and network policy primitives. KubeVirt reached 1.0 in 2023 and is the engine behind Red Hat OpenShift Virtualization, the downstream product used by organisations migrating away from VMware.