Skip to main content
  1. Index/

Inference

Inference is the operational phase of machine learning where a trained model is applied to new inputs to produce outputs: next tokens in an LLM, bounding boxes in vision, embeddings for search, or scores in tabular models. Its objective is reliable serving at scale—honoring latency targets (time to first token, p99 completion time), throughput (requests or tokens per second), availability, and cost per query—rather than improving weights. In generative AI, inference splits into prefill (processing the prompt in one or few forward passes) and decode (autoregressive generation of each output token), each with different bottlenecks. Production inference adds API gateways, auth, rate limiting, observability, model versioning, A/B tests, and guardrails; the model file is read-mostly while KV cache and batch state are ephemeral per session.

Architecturally, inference favors low latency per request and efficient memory reuse on accelerators; training favors sustained FLOPs over huge datasets with frequent all-to-all communication and checkpoint writes. A CPU can serve small models or orchestrate I/O, but LLM inference at useful scale runs on GPUs (or TPUs/ASICs) with frameworks such as vLLM, NIM, or Triton, often in containers on Kubernetes. The host CPU schedules batches, handles networking, and runs the control plane; the device holds weights and the KV cache. Scaling is horizontal (more replicas) and vertical (tensor parallelism, disaggregation); schedulers like llm-d route traffic so replicas with warm prefix cache do more work. Training clusters optimize for gradient sync and large memory per job; inference clusters optimize for concurrent sessions, tail latency, and cost per token—different failure modes and different autoscaling signals.

Red Hat addresses inference through Red Hat OpenShift AI, RHEL AI, and OpenShift as the deployment substrate. Supported patterns include vLLM and partner NIM microservices on GPU nodes, routes and service mesh for north-south traffic, llm-d for distributed routing on OpenShift, and enterprise Linux for drivers and security. Red Hat documents reference architectures for private AI inference (GPU sizing, MIG, networking, storage for model artifacts), integrates observability and CI/CD for model promotions, and aligns with the NVIDIA AI stack on certified hardware. Training may happen elsewhere (cloud burst, dedicated cluster); inference is often what runs continuously on the customer’s Red Hat platform close to applications and data.

Related

Quantization

Quantization is the process of representing a model’s weights and/or activations with fewer bits than full FP32 training precision—commonly FP16, BF16, FP8, INT8, or INT4 (GPTQ, AWQ, GGUF-style formats). The objective is lower GPU memory (larger models or more concurrent sessions per card), higher throughput, and sometimes faster kernels on hardware with native low-precision units, at the cost of possible quality degradation if pushed too aggressively. Quantization can be applied post-training (calibration on a sample dataset) or during training (quantization-aware training). For inference, serving engines vLLM and NIM load quantized checkpoints and dispatch to vendor libraries (TensorRT-LLM, CUTLASS, etc.) that implement fused low-precision matmuls.

vLLM

vLLM is an open-source library and serving stack for large language model (LLM) inference. Its objective is to turn a trained model into a production service that sustains many concurrent users with low latency and high tokens per second per GPU. vLLM targets the inference phase (prefill + decode), not training: it loads weights onto accelerators, batches incoming prompts, schedules decode steps, and streams completions back to clients over HTTP/gRPC (often via an OpenAI-compatible API). It has become a de facto engine behind many private and cloud AI gateways because it ships integrations for Hugging Face models, LoRA adapters, tensor parallelism, pipeline parallelism, speculative decoding, and quantization (GPTQ, AWQ, FP8).

Decode

Decode is the second stage of LLM inference: after prefill has stored keys and values for the prompt, the model generates one new token per forward pass, appends it to the sequence, extends the KV cache, and repeats until a stop condition (EOS token, max length, or API limit). Its objective is fluent continuation—answer text, code, or tool-call JSON—at acceptable inter-token latency and cluster throughput (tokens per second across many concurrent sessions). Decode drives the “typing” experience in chat UIs; prefill drives how long users wait before the first character appears.