Skip to main content

GPU Sharing & Multi-tenancy

Your GPU cluster is sitting mostly idle. Teams wait days for GPU access. Training jobs fail because the scheduler doesn't understand GPU topology. Kubernetes can run GPUs well. It just needs someone who has done it before.

Utilization, not idle silicon

MIG partitioning, time-slicing and a GPU-aware scheduler so more teams share the same hardware without stepping on each other.

We know the scheduler code

We are active contributors to KAI Scheduler, so gang scheduling, queues and preemption are things we tune from the inside, not from the docs.

Any substrate

EKS with EFA, GKE GPU pools with multi-networking, or bare metal with Calico/Cilium. We have shipped all three.

What we do

Make one cluster serve many teams.

The operator stack, the scheduler and the tenancy model, configured so GPUs are scheduled, shared and accounted for.

01
GPU Operator stack
Driver containers, Container Toolkit, Device Plugin, DCGM Exporter and GPU Feature Discovery. We handle driver conflicts, runtime differences, secure boot and upgrade rollouts.
02
KAI Scheduler
Topology-aware placement, fair-share scheduling, gang scheduling for distributed training, preemption policies and queue management. We are an active contributor.
03
MIG & fractional sharing
A100 and H100 MIG partitioning, time-slicing for non-MIG GPUs, and workload-aware partition profiles.
04
Multi-tenancy
Namespace isolation, GPU resource quotas, RBAC, priority classes, and cost allocation and chargeback.
05
Job orchestration
Kubeflow Training Operator and PyTorchJob, with integration into MLflow and W&B for training workflows.
How we work

Scope, design, build, hand off.

01

Assess

Audit your Kubernetes GPU setup, scheduler config and utilization.

02

Design

Right-size the operator stack, scheduling policy and tenancy model.

03

Implement

Deploy, configure and validate with real workloads.

04

Transfer

Runbooks, dashboards and training for your platform team.

Our engineers contribute upstream to the projects this layer runs on: KAI Scheduler, Network Operator, DOCA driver build, ipoib-cni. See the contributions →
Technologies we work with
KubernetesGPU OperatorKAI SchedulerNetwork OperatorMIGTime-SlicingEKSGKEKubeflowMultusSR-IOVHelmArgoCD
FAQ
What is the NVIDIA GPU Operator?
A set of Kubernetes operators that automate the lifecycle of GPU drivers, Container Toolkit, device plugin, DCGM exporter and MIG manager across every GPU node. You need it any time you want GPUs scheduled as Kubernetes resources.
How does GPU sharing work in Kubernetes?
There are three modes: MIG for hardware partitioning on A100 and H100 (hard isolation, fixed sizes), time-slicing for simple time-multiplexing (no isolation), and MPS for CUDA-level process sharing. Pick MIG for multi-tenant production and time-slicing for dev and inference.
What is the KAI Scheduler?
KAI Scheduler (formerly the Run:ai scheduler) is a Kubernetes-native gang scheduler purpose-built for GPU workloads: queues, fair-share, gang scheduling and preemption with GPU-awareness. We are an active contributor to this project.
Can I run GPU workloads on EKS, GKE, or AKS?
Yes. All three support GPU node groups and the GPU Operator runs on top. The complications are around driver versions, instance-type-specific CUDA images, multi-tenancy isolation and in-cluster networking (especially RDMA or EFA).
Do I need Slurm if I already run Kubernetes?
Not usually. Kubernetes with GPU Operator, KAI or Volcano scheduler, and the Kubeflow Training Operator covers most distributed-training workloads. Slurm still wins for traditional HPC or organizations with deep Slurm operational expertise.
How do you approach multi-tenant GPU clusters?
Namespace quotas, ResourceQuotas on nvidia.com/gpu, a GPU-aware scheduler for fair-share, MIG or SR-IOV for hardware isolation where needed, node taints and tolerations for workload separation, and per-namespace DCGM metrics for visibility.
Need help running GPUs on Kubernetes?

We've deployed GPU Operator and KAI Scheduler on EKS, GKE and bare metal. Let's look at your cluster.