Skip to main content

AI Factory Setup

You're building a GPU cluster on hardware that's been delivered, on-prem, colo, or dedicated cloud. You want compute, networking, storage, orchestration, and monitoring right the first time, without spending months figuring out what NVIDIA's docs don't tell you.

Full-stack architecture

Every layer from node bring-up to the scheduler, designed as one system with a written architecture document.

Production-ready on day 1

The first training job runs on a cluster with verified RDMA, working scheduling, monitoring and runbooks, not a lab.

We do the implementation

Provisioning, configuration and testing on the installed hardware, by the engineers you scoped it with.

What we do

Five layers, one cluster.

We work on everything above the installed hardware. Rack, power and cooling belong to you or your data-centre partner.

01
Node bring-up
BMC discovery and inventory, automated OS provisioning, driver / CUDA / Fabric Manager stack, firmware baselines, NVLink and NVSwitch topology verification, node build consistency checks.
02
Network fabric
RDMA fabric configuration (InfiniBand or RoCE) with PFC/ECN, compute/storage network separation, VXLAN/VRF on the switches where needed, and GPUDirect RDMA verified across the fabric.
03
Storage integration
Parallel filesystem integration (Lustre, WekaFS, GPFS), checkpoint paths, data staging, GPUDirect Storage where the workload benefits.
04
Orchestration & scheduling
Kubernetes with GPU Operator and KAI Scheduler, or Slurm with Pyxis/Enroot. Multi-tenancy, quotas, gang scheduling, job queues.
05
Operations
DCGM monitoring, XID error detection, automated fault recovery, capacity planning, upgrade procedures, runbooks and handover to your team.
How we work

Scope, design, build, hand off.

01

Scope

Understand your workload, hardware, timeline and who operates the cluster afterwards.

02

Design

Architecture document covering all five layers, with the configuration decisions written down.

03

Build

Provision, configure, verify. We do the implementation on the installed hardware.

04

Hand off

Runbooks, dashboards, and knowledge transfer so your team can run it.

Our engineers contribute upstream to the projects this layer runs on: KAI Scheduler, Network Operator, DOCA driver build, ipoib-cni. See the contributions →
Technologies we work with
H100H200B200GH200MI300XDGX / HGXInfiniBandRoCESpectrum-XKubernetesSlurmGPU OperatorNetwork OperatorKAI SchedulerLustreWekaFSDCGMPrometheusGrafana
FAQ
What is an AI factory?
A full-stack GPU compute environment purpose-built for AI training and inference (compute, high-speed networking, storage, orchestration, observability, and tenancy) operated as a product for internal or external AI teams.
How long does it take to set up a production GPU cluster?
For a well-scoped deployment on dedicated hardware, a functional training-ready GPU cluster is typically weeks, not months.
Should I build on-prem, in a colo, or in the cloud?
Cloud is fastest to start and best for bursty workloads. Colo and on-prem win on cost once utilization is consistently high. We'll tell you which on the scoping call; we have no stake in the answer.
What storage architecture do I need?
Training I/O is dominated by large-file sequential reads and checkpoint writes. We size the storage path from your dataset, checkpoint frequency and model size, then integrate the filesystem you already have or help you pick one.
How do you size the network fabric?
We size inter-node bandwidth from the model's gradient volume and target AllReduce-to-compute ratio, then verify the configured fabric with nccl-tests before the first real job.
Do you operate the cluster after it is built?
Both. We lead new-cluster builds and can hand off to your SRE/platform team with documentation and runbooks, or stay on for Day-2 operations.
Planning a GPU cluster build?

We've done this before. Let's talk about what you're building.