C
CYSKA
Consulting
EN FR
Navigation
Sovereign AI · Opportunity study

The sovereign platform underneath your AI models

Running open-weight models, fine-tuning them on internal data or giving your data teams private GPU capacity raises the same questions as any sovereign cloud: where the data and the model weights live, who can access them, what it costs at steady state, and how you would leave. OpenStack answers them with GPU compute, storage and isolation you own. We help you decide whether it is worth it for you, and design it if it is.

What sits where
Layers
Your teams
Models, data pipelines, evaluation, applications (vLLM, PyTorch, notebooks, RAG)
Orchestration
Kubernetes with the GPU Operator, or Slurm for batch training
Our scope: OpenStack
GPU instances (passthrough, vGPU, MIG) or bare-metal nodes, Ceph storage for datasets and checkpoints, per-project isolation, quotas and reservations
Hardware
GPU servers, fast network, power and cooling, in your datacenter or a sovereign colocation
⚠

Where we stand, honestly

We have not yet operated a GPU cluster on OpenStack in production for a client, so we publish no field results on this topic. What we bring today is the opportunity study, the sizing and the design, built on our production practice of OpenStack, Ceph and Kolla-Ansible and on the official documentation. The first GPU pilot is scoped as a pilot, with measured figures, before anyone commits to a programme.

Use cases

Three realistic reasons to bring AI workloads in-house

Not every AI workload belongs on your own GPUs. These three do, when data sensitivity or steady usage makes the public cloud the wrong default.

01

Inference of open-weight models

Serve Llama, Mistral, Qwen or domain models behind your own API for assistants, document processing or code generation. Prompts and outputs never leave your perimeter; costs are fixed instead of per-token.

02

Fine-tuning on internal data

Adapt a base model to your vocabulary, procedures or customer history. The training set is often the most sensitive asset you own; keeping it and the resulting weights on infrastructure you control is a governance decision, not a technical preference.

03

Private GPU capacity for data teams

A self-service pool of GPU instances with projects, quotas and reservations, so research and product teams stop competing for a shared workstation or filing cloud purchase requests for every experiment.

The decision

Public cloud GPUs or your own? An honest checklist

The opportunity study answers each line with your numbers. If the answer is “stay in the cloud”, we say so.

Criterion Favours your own OpenStack GPUs Favours the public cloud
Usage profile Steady inference or recurring fine-tuning, GPUs busy most of the time Bursty experiments, a few large training runs per year
Data sensitivity Regulated, classified or contractually confined data; model weights considered a trade secret Public or synthetic data, no residency constraint
Model size Models that fit on one to a few GPUs (7B to 70B class, quantized when needed) Frontier-scale pre-training needing hundreds of interconnected GPUs
Existing platform An OpenStack and Ceph platform already operated in-house: GPUs become one more compute class No private cloud and no team to operate one
Facilities Datacenter or sovereign colocation with the power and cooling GPU servers require No suitable room, or a hardware lead time you cannot absorb
Reversibility You want to be able to change hardware vendor, model provider or hosting without rewriting the stack Dependency on one provider's managed AI services is acceptable
Why OpenStack

GPUs become one more resource of the platform you already govern

The same projects, quotas, identities, networks and storage that serve your virtual machines serve your AI workloads. No second platform, no second team, no second audit scope.

Technical guide: GPU passthrough and vGPU with Nova →

Three ways to expose a GPU

Whole card to one instance (passthrough), shared slices with vendor drivers (vGPU or MIG), or a bare-metal node when the framework needs the hardware directly.

Storage built for datasets

Ceph block volumes for checkpoints, shared file systems for training sets, object storage for model registries, all under the same quotas and encryption policies.

Kubernetes where the tools expect it

Clusters provisioned on GPU instances, with the GPU Operator for scheduling and vLLM, Ray or Kubeflow on top. Your data teams keep the tooling they know.

Isolation and reservations

One project per team or per sensitivity level, GPU quotas, time-boxed reservations for training campaigns, and the same observability chain as the rest of the platform.

What nobody puts on the slide

The hardware realities the study puts a number on

Cost and lead time of GPUs

Datacenter GPUs are expensive and supply-constrained. The study compares total cost over three to five years with your actual cloud bill, not with list prices.

Power and cooling

A GPU server draws several kilowatts. Rack density, cooling and electrical capacity are validated with your facilities team before any hardware is ordered.

Driver licensing

Giving a whole GPU to one instance is free; sharing a card between instances with the vendor's vGPU software is a paid licence. The choice changes both the budget and the operating model.

Network for multi-GPU training

Training across several nodes needs low-latency, high-bandwidth interconnects. Inference and single-node fine-tuning usually do not. Sizing the network to the real workload avoids the most common over-spend.

01

Opportunity study

Inventory of AI workloads and their data, usage profile, current cloud spend, regulatory constraints, facilities check. Output: a go / no-go recommendation with your numbers, including “stay in the cloud” when that is the honest answer.

02

Design and sizing

GPU exposure model, flavors and quotas, storage layout for datasets and checkpoints, network, Kubernetes integration, observability, security and reversibility plan. Hardware bill of materials validated with your suppliers.

03

Measured pilot

One or two GPU nodes on your OpenStack, one real inference or fine-tuning workload, measured throughput, cost and operability. Only then is a programme sized. That pilot is where our first field figures on this topic will come from.

What your team receives

An opportunity report with a clear recommendation, a target architecture, a sizing and cost model over several years, a hardware and licensing checklist, a pilot plan with acceptance criteria, and the reversibility plan written before the first GPU is ordered.

FAQ

Frequently asked questions

Do you build or train the models themselves?

No. Data science, model selection and evaluation stay with your teams or your AI partner. Our scope is the platform underneath: GPU compute, storage, isolation, Kubernetes integration and operations on OpenStack.

Can we start small?

That is the only way we recommend. One or two GPU nodes added to an existing OpenStack, one real workload, measured results. The platform is designed so that adding nodes later is a procurement decision, not a redesign.

We do not run OpenStack today. Does this still apply?

Yes, but the study will weigh it honestly: a GPU platform is rarely a good reason on its own to adopt a private cloud. If you are also leaving VMware or consolidating infrastructure, the two decisions reinforce each other.

Which GPUs and which vendor?

The study is vendor-neutral: the choice follows the workload (memory per model, sharing needs, framework support) and availability, not a partnership. Passthrough works with any PCI GPU; sharing a card requires the vendor's virtualization software and its licence.