Skip to main content
  1. Index/

CUDA (Compute Unified Device Architecture)

CUDA (Compute Unified Device Architecture) is NVIDIA’s software platform for general-purpose computing on GPUs. Its objective is to give developers a familiar C/C++ (and Fortran, Python bindings) programming model with explicit control over kernels (functions that run on the device), streams (ordered queues of work), and memory spaces (host, device, unified). CUDA sits above the GPU driver and below frameworks such as cuDNN, NCCL, and higher-level ML stacks; it is the layer that makes it practical to implement custom operators, HPC solvers, and inference engines that are not covered by a closed library. For AI, virtually every major training and inference stack ultimately depends on CUDA (or a CUDA-compatible runtime) on NVIDIA hardware.

The CUDA execution model is SIMT (Single Instruction, Multiple Threads): programmers write kernels as scalar thread programs; the hardware groups threads into warps of 32 that execute in lockstep. Global memory is large but high-latency; shared memory and registers are fast but limited per block. A CPU thread is heavyweight, preemptible, and optimized for irregular control flow and system calls; a CUDA thread is extremely lightweight and only efficient when workloads are regular and memory access is coalesced. Host code runs on the CPU and issues asynchronous launches; synchronization (cudaDeviceSynchronize, events) defines visibility between host and device. Multi-GPU scaling uses peer access, NCCL collectives, and NVLink/PCIe topology awareness—concerns that do not exist in single-socket CPU programming in the same form.

Red Hat does not ship CUDA itself (it is NVIDIA proprietary) but documents and supports running CUDA workloads on RHEL and OpenShift. RHEL notes cover installing the NVIDIA driver, CUDA toolkit from NVIDIA repositories or containers, and using the NVIDIA Container Toolkit so GPU devices pass through to Podman or Kubernetes pods. OpenShift AI and OpenShift GPU operator patterns integrate device plugins and validated images for notebook and model-serving workloads. Red Hat’s AI portfolio assumes CUDA where NVIDIA GPUs are present, while keeping the operating platform (SELinux, cgroups, kABI-stable drivers where offered) under enterprise support boundaries defined in Red Hat and NVIDIA joint certification guides.

Related

NCCL (NVIDIA Collective Communications Library)

NCCL (NVIDIA Collective Communications Library) implements collective operations—all-reduce, broadcast, reduce-scatter, all-gather, and others—optimized for NVIDIA GPUs across NVLink within a node and RDMA (InfiniBand or RoCE) across nodes. Its objective in AI is to make distributed training and multi-GPU inference (tensor parallelism) scale: gradient shards must merge every step; attention and MLP partitions must exchange activations with minimal latency. Frameworks (PyTorch DDP/FSDP, vLLM tensor parallel) call NCCL (or delegate to it) rather than hand-rolling socket code.

cuDNN (CUDA Deep Neural Network library)

cuDNN (CUDA Deep Neural Network library) is NVIDIA’s library of highly optimized GPU kernels for operations that dominate deep learning: convolutions, matrix multiplies used in attention, pooling, normalization (batch/layer), activations, and recurrent cells. Its objective is to deliver near-peak performance on CUDA-capable GPUs without every framework author hand-writing assembly-tuned kernels. PyTorch, TensorFlow, and many inference engines call cuDNN (directly or via cuBLAS) under the hood for training and serving. cuDNN sits between raw CUDA and application code; version alignment with the CUDA toolkit and driver is mandatory for supported deployments.

Confidential GPU

A Confidential GPU is a GPU whose memory, computation state, and data transfers are hardware-encrypted and isolated from the host system — extending the Trusted Execution Environment (TEE) boundary that technologies like TDX and SEV-SNP provide at the CPU level to encompass the GPU accelerator as well. The primary implementation today is NVIDIA Confidential Computing on the Hopper architecture (H100 and later), which encrypts all data resident in GPU High Bandwidth Memory (HBM) using per-context keys managed by the GPU’s on-die security processor. This means that model weights, training data, activations, and intermediate computations are cryptographically protected throughout GPU processing — a host administrator, hypervisor, or co-tenant with DMA access to the PCIe bus sees only ciphertext. The GPU also participates in a dedicated attestation flow: the NVIDIA Remote Attestation Service (NRAS) produces signed evidence that a specific GPU is genuine NVIDIA hardware running in Confidential Computing mode with unmodified firmware, analogous to how Intel DCAP or AMD KDS attest CPU TEEs. This GPU attestation is verified alongside CPU attestation before secrets (model decryption keys, dataset credentials) are released to the combined CPU+GPU TEE. The technology requires no application code changes — existing TensorFlow, PyTorch, and CUDA workloads run unmodified inside the confidential boundary. The primary threat model is the same as CPU-level confidential computing (protecting data-in-use from the infrastructure operator) but applied to the specific risk of AI workloads: model intellectual property theft, training data exfiltration, and inference input/output interception during GPU computation.