The sovereign platform underneath your AI models
Running open-weight models, fine-tuning them on internal data or giving your data teams private GPU capacity raises the same questions as any sovereign cloud: where the data and the model weights live, who can access them, what it costs at steady state, and how you would leave. OpenStack answers them with GPU compute, storage and isolation you own. We help you decide whether it is worth it for you, and design it if it is.
Where we stand, honestly
We have not yet operated a GPU cluster on OpenStack in production for a client, so we publish no field results on this topic. What we bring today is the opportunity study, the sizing and the design, built on our production practice of OpenStack, Ceph and Kolla-Ansible and on the official documentation. The first GPU pilot is scoped as a pilot, with measured figures, before anyone commits to a programme.
Three realistic reasons to bring AI workloads in-house
Not every AI workload belongs on your own GPUs. These three do, when data sensitivity or steady usage makes the public cloud the wrong default.
Inference of open-weight models
Serve Llama, Mistral, Qwen or domain models behind your own API for assistants, document processing or code generation. Prompts and outputs never leave your perimeter; costs are fixed instead of per-token.
Fine-tuning on internal data
Adapt a base model to your vocabulary, procedures or customer history. The training set is often the most sensitive asset you own; keeping it and the resulting weights on infrastructure you control is a governance decision, not a technical preference.
Private GPU capacity for data teams
A self-service pool of GPU instances with projects, quotas and reservations, so research and product teams stop competing for a shared workstation or filing cloud purchase requests for every experiment.
Public cloud GPUs or your own? An honest checklist
The opportunity study answers each line with your numbers. If the answer is “stay in the cloud”, we say so.
| Criterion | Favours your own OpenStack GPUs | Favours the public cloud |
|---|---|---|
| Usage profile | Steady inference or recurring fine-tuning, GPUs busy most of the time | Bursty experiments, a few large training runs per year |
| Data sensitivity | Regulated, classified or contractually confined data; model weights considered a trade secret | Public or synthetic data, no residency constraint |
| Model size | Models that fit on one to a few GPUs (7B to 70B class, quantized when needed) | Frontier-scale pre-training needing hundreds of interconnected GPUs |
| Existing platform | An OpenStack and Ceph platform already operated in-house: GPUs become one more compute class | No private cloud and no team to operate one |
| Facilities | Datacenter or sovereign colocation with the power and cooling GPU servers require | No suitable room, or a hardware lead time you cannot absorb |
| Reversibility | You want to be able to change hardware vendor, model provider or hosting without rewriting the stack | Dependency on one provider's managed AI services is acceptable |
GPUs become one more resource of the platform you already govern
The same projects, quotas, identities, networks and storage that serve your virtual machines serve your AI workloads. No second platform, no second team, no second audit scope.
Technical guide: GPU passthrough and vGPU with Nova →Three ways to expose a GPU
Whole card to one instance (passthrough), shared slices with vendor drivers (vGPU or MIG), or a bare-metal node when the framework needs the hardware directly.
Storage built for datasets
Ceph block volumes for checkpoints, shared file systems for training sets, object storage for model registries, all under the same quotas and encryption policies.
Kubernetes where the tools expect it
Clusters provisioned on GPU instances, with the GPU Operator for scheduling and vLLM, Ray or Kubeflow on top. Your data teams keep the tooling they know.
Isolation and reservations
One project per team or per sensitivity level, GPU quotas, time-boxed reservations for training campaigns, and the same observability chain as the rest of the platform.
The hardware realities the study puts a number on
Cost and lead time of GPUs
Datacenter GPUs are expensive and supply-constrained. The study compares total cost over three to five years with your actual cloud bill, not with list prices.
Power and cooling
A GPU server draws several kilowatts. Rack density, cooling and electrical capacity are validated with your facilities team before any hardware is ordered.
Driver licensing
Giving a whole GPU to one instance is free; sharing a card between instances with the vendor's vGPU software is a paid licence. The choice changes both the budget and the operating model.
Network for multi-GPU training
Training across several nodes needs low-latency, high-bandwidth interconnects. Inference and single-node fine-tuning usually do not. Sizing the network to the real workload avoids the most common over-spend.
Opportunity study
Inventory of AI workloads and their data, usage profile, current cloud spend, regulatory constraints, facilities check. Output: a go / no-go recommendation with your numbers, including “stay in the cloud” when that is the honest answer.
Design and sizing
GPU exposure model, flavors and quotas, storage layout for datasets and checkpoints, network, Kubernetes integration, observability, security and reversibility plan. Hardware bill of materials validated with your suppliers.
Measured pilot
One or two GPU nodes on your OpenStack, one real inference or fine-tuning workload, measured throughput, cost and operability. Only then is a programme sized. That pilot is where our first field figures on this topic will come from.
What your team receives
An opportunity report with a clear recommendation, a target architecture, a sizing and cost model over several years, a hardware and licensing checklist, a pilot plan with acceptance criteria, and the reversibility plan written before the first GPU is ordered.
Frequently asked questions
Do you build or train the models themselves?
No. Data science, model selection and evaluation stay with your teams or your AI partner. Our scope is the platform underneath: GPU compute, storage, isolation, Kubernetes integration and operations on OpenStack.
Can we start small?
That is the only way we recommend. One or two GPU nodes added to an existing OpenStack, one real workload, measured results. The platform is designed so that adding nodes later is a procurement decision, not a redesign.
We do not run OpenStack today. Does this still apply?
Yes, but the study will weigh it honestly: a GPU platform is rarely a good reason on its own to adopt a private cloud. If you are also leaving VMware or consolidating infrastructure, the two decisions reinforce each other.
Which GPUs and which vendor?
The study is vendor-neutral: the choice follows the workload (memory per model, sharing needs, framework support) and availability, not a partnership. Passthrough works with any PCI GPU; sharing a card requires the vendor's virtualization software and its licence.