Skip to main content
  1. Index/

vLLM

Table of Contents

vLLM is an open-source library and serving stack for large language model (LLM) inference. Its objective is to turn a trained model into a production service that sustains many concurrent users with low latency and high tokens per second per GPU. vLLM targets the inference phase (prefill + decode), not training: it loads weights onto accelerators, batches incoming prompts, schedules decode steps, and streams completions back to clients over HTTP/gRPC (often via an OpenAI-compatible API). It has become a de facto engine behind many private and cloud AI gateways because it ships integrations for Hugging Face models, LoRA adapters, tensor parallelism, pipeline parallelism, speculative decoding, and quantization (GPTQ, AWQ, FP8).

Architecturally, vLLM differs from running a model naively on a CPU or a single-threaded GPU script in three ways. First, PagedAttention stores the KV cache in non-contiguous GPU memory pages (like virtual memory), so batch size and sequence length scale without the fragmentation that wastes VRAM on fixed buffers. Second, continuous batching adds and removes sequences from a running batch each iteration instead of waiting for every request in a static batch to finish—critical for chat workloads with uneven lengths. Third, custom CUDA kernels and communication backends (NCCL) keep attention and MLP matmuls on the device while the host CPU only orchestrates scheduling and I/O. A CPU-only path exists for tiny models but is not the design center; vLLM assumes GPU (or supported accelerator) memory bandwidth and parallelism dominate cost and latency.

Red Hat integrates vLLM into Red Hat OpenShift AI and RHEL AI reference architectures as a supported model-serving option alongside other runtimes. On OpenShift, vLLM runs in containers with GPU device plugins, often behind a route or gateway; Red Hat documents validated GPU nodes, drivers, and partner stacks (including NVIDIA). llm-d, a Kubernetes-native inference project co-led by Red Hat and others, uses vLLM as its default model server and adds cluster-level routing and disaggregation above it. Operators therefore deploy vLLM as the per-pod engine while Red Hat platforms supply Linux, Kubernetes, security (SELinux, SCCs), and MLOps workflows around it.

Additional Information
#

Related

Inference

Inference is the operational phase of machine learning where a trained model is applied to new inputs to produce outputs: next tokens in an LLM, bounding boxes in vision, embeddings for search, or scores in tabular models. Its objective is reliable serving at scale—honoring latency targets (time to first token, p99 completion time), throughput (requests or tokens per second), availability, and cost per query—rather than improving weights. In generative AI, inference splits into prefill (processing the prompt in one or few forward passes) and decode (autoregressive generation of each output token), each with different bottlenecks. Production inference adds API gateways, auth, rate limiting, observability, model versioning, A/B tests, and guardrails; the model file is read-mostly while KV cache and batch state are ephemeral per session.

Quantization

Quantization is the process of representing a model’s weights and/or activations with fewer bits than full FP32 training precision—commonly FP16, BF16, FP8, INT8, or INT4 (GPTQ, AWQ, GGUF-style formats). The objective is lower GPU memory (larger models or more concurrent sessions per card), higher throughput, and sometimes faster kernels on hardware with native low-precision units, at the cost of possible quality degradation if pushed too aggressively. Quantization can be applied post-training (calibration on a sample dataset) or during training (quantization-aware training). For inference, serving engines vLLM and NIM load quantized checkpoints and dispatch to vendor libraries (TensorRT-LLM, CUTLASS, etc.) that implement fused low-precision matmuls.

Decode

Decode is the second stage of LLM inference: after prefill has stored keys and values for the prompt, the model generates one new token per forward pass, appends it to the sequence, extends the KV cache, and repeats until a stop condition (EOS token, max length, or API limit). Its objective is fluent continuation—answer text, code, or tool-call JSON—at acceptable inter-token latency and cluster throughput (tokens per second across many concurrent sessions). Decode drives the “typing” experience in chat UIs; prefill drives how long users wait before the first character appears.