C
CYSKA
Consulting
EN FR
Navigation
OpenStack · GPU · Design guide

Expose GPUs to OpenStack instances: PCI passthrough and vGPU with Nova and Kolla-Ansible

The full chain to give a GPU to an instance: host preparation (IOMMU, vfio-pci), Nova configuration through Kolla-Ansible (device_spec, alias, scheduler filter), GPU flavors and image properties, then the shared alternative with vGPU or MIG through mediated devices. Ends with Placement checks and the pitfalls to test on a pilot.

Audience: platform engineers Reference: Nova 2024.1+, Kolla-Ansible, KVM, NVIDIA datacenter GPUs Updated: September 2026

Status of this guide

This is a design guide, not a field report: the configuration below is taken from the official Nova, Kolla-Ansible and NVIDIA documentation and cross-checked, but we have not yet run it in production for a client. Treat every block as a starting point to validate on a pilot node before any production commitment. We will update this page with measured results once we have them.

1. Three ways to expose a GPU

PCI passthrough

The whole card is handed to one instance through VFIO. Native performance, no vendor licence, works with any PCI GPU. One card, one instance; no live migration by default.

vGPU / MIG (mediated devices)

The host driver slices a card into several virtual GPUs exposed as mdev devices; Nova schedules them as VGPU resources. Requires the vendor's virtualization software (NVIDIA vGPU is licensed). Ideal for many small inference or notebook workloads.

Bare metal (Ironic)

The whole server is provisioned by Ironic and the GPUs are used directly by the OS or a Kubernetes node. No virtualization overhead, full NVLink/RDMA access; the least flexible option. Not covered further here.

2. Prepare the compute host

Passthrough needs VT-d or AMD-Vi enabled in the firmware, the IOMMU enabled in the kernel, and the GPU bound to vfio-pci instead of the vendor driver. Identify the vendor and product IDs first: you will reuse them in every Nova option. Bind the audio function of the card too when it shares the IOMMU group.

# 1. Identify the GPU: PCI address and vendor:product IDs (NVIDIA vendor id is 10de)
lspci -nn | grep -i nvidia
#   41:00.0 3D controller [0302]: NVIDIA Corporation ... [10de:2331]

# 2. Enable the IOMMU (Intel: intel_iommu=on ; AMD: amd_iommu=on), then reboot
sed -i 's/GRUB_CMDLINE_LINUX="/GRUB_CMDLINE_LINUX="intel_iommu=on iommu=pt /' /etc/default/grub
update-grub && reboot
dmesg | grep -i -e DMAR -e IOMMU | head          # after reboot: "IOMMU enabled"

# 3. Bind the GPU to vfio-pci and keep the vendor driver off the host
cat > /etc/modprobe.d/vfio.conf <<'EOF'
options vfio-pci ids=10de:2331
softdep nvidia pre: vfio-pci
softdep nouveau pre: vfio-pci
EOF
echo "blacklist nouveau" > /etc/modprobe.d/blacklist-nouveau.conf
echo vfio-pci > /etc/modules-load.d/vfio-pci.conf
update-initramfs -u && reboot

# 4. Check the binding and the IOMMU group
lspci -nnk -s 41:00.0 | grep -i 'kernel driver'   # Kernel driver in use: vfio-pci
find /sys/kernel/iommu_groups/ -type l | grep 41:00

3. Configure Nova through Kolla-Ansible

Three services need to know about the GPU: nova-compute (which devices may be passed through, via device_spec), nova-api and nova-scheduler (the alias users request in flavors, and the PciPassthroughFilter). Kolla-Ansible merges per-service overrides from /etc/kolla/config/nova/, and per-host overrides from a sub-directory named after the inventory host, so GPU settings only reach GPU nodes.

# /etc/kolla/config/nova/gpu-node-01/nova.conf   (per-host override: only the GPU computes)
[pci]
# Every device matching vendor/product becomes assignable; add "address" to pin specific slots
device_spec = { "vendor_id": "10de", "product_id": "2331" }
# Optional since Zed: publish PCI inventory to Placement (rollback is not supported once enabled on a host)
report_in_placement = true

# /etc/kolla/config/nova/nova-api.conf   and   /etc/kolla/config/nova/nova-scheduler.conf
[pci]
alias = { "vendor_id": "10de", "product_id": "2331", "device_type": "type-PCI", "name": "h100", "numa_policy": "preferred" }

[filter_scheduler]
enabled_filters = ComputeFilter,ComputeCapabilitiesFilter,ImagePropertiesFilter,ServerGroupAntiAffinityFilter,ServerGroupAffinityFilter,PciPassthroughFilter,AggregateInstanceExtraSpecsFilter
available_filters = nova.scheduler.filters.all_filters
pci_in_placement = true

# The alias must also be visible to nova-compute: put the same [pci] alias line in the per-host file above.

# Apply with Kolla-Ansible (nova only), then check the compute picked up the devices
kolla-ansible -i ./multinode reconfigure --tags nova
docker logs nova_compute 2>&1 | grep -i pci | tail

Place GPU nodes in a dedicated host aggregate and bind GPU flavors to its metadata with AggregateInstanceExtraSpecsFilter. This directs GPU flavors to those hosts, but does not by itself prevent ordinary flavors from landing there. If the nodes must be reserved, add and test an explicit scheduling or access policy.

openstack aggregate create --zone GPU gpu
openstack aggregate set --property gpu=true gpu
openstack aggregate add host gpu gpu-node-01

4. GPU flavor, image and first boot

The flavor requests the alias (name:count). Datacenter GPUs expose very large PCI BARs: boot the guest as a q35 machine with UEFI firmware, otherwise the device may not initialise. The guest then needs the vendor driver and, for containers, the NVIDIA container toolkit.

# Flavor: 1 GPU via the alias, pinned to the GPU aggregate
openstack flavor create --vcpus 16 --ram 131072 --disk 200 g1.h100.1
openstack flavor set g1.h100.1 \
  --property "pci_passthrough:alias"="h100:1" \
  --property "aggregate_instance_extra_specs:gpu"="true" \
  --property "hw:pci_numa_affinity_policy"="preferred"

# Image: q35 + UEFI so large-BAR GPUs initialise correctly
openstack image set --property hw_machine_type=q35 --property hw_firmware_type=uefi ubuntu-24.04-gpu

# Boot and verify from inside the guest
openstack server create --flavor g1.h100.1 --image ubuntu-24.04-gpu --network <NET_ID> --key-name <KEY> gpu-test-01
#   guest$ lspci -nn | grep -i nvidia
#   guest$ install the exact driver validated by the NVIDIA compatibility matrix, then reboot
#   guest$ nvidia-smi

5. Sharing a card: vGPU and MIG

For many small workloads, one card per instance wastes capacity. With the NVIDIA vGPU host driver (licensed) the card is exposed as mediated devices of a given type, and Nova schedules them as VGPU resources. On Ampere and later, MIG partitions the GPU in hardware first; each MIG instance is then exposed as an mdev type. In both cases the host driver, not vfio-pci, must own the card, so this configuration excludes passthrough on the same device.

# Host: NVIDIA vGPU host driver installed (licensed), card NOT bound to vfio-pci
nvidia-smi                                       # host driver sees the card
ls /sys/class/mdev_bus/                          # PCI addresses that can create mediated devices
ls /sys/class/mdev_bus/0000:41:00.0/mdev_supported_types/
cat /sys/class/mdev_bus/0000:41:00.0/mdev_supported_types/nvidia-*/name   # e.g. GRID H100-20C

# Optional, Ampere+: MIG partitions (hardware slices) before vGPU
nvidia-smi -i 0 -mig 1
nvidia-smi mig -i 0 -cgi 9,9,9 -C                # three 3g.40gb instances on an 80 GB card, profile ids vary per model
/usr/lib/nvidia/sriov-manage -e 0000:41:00.0     # enable the SR-IOV VFs that carry the MIG-backed vGPUs

# /etc/kolla/config/nova/gpu-node-02/nova.conf
[devices]
enabled_mdev_types = nvidia-XXX                  # the mdev type name from mdev_supported_types

[mdev_nvidia-XXX]
device_addresses = 0000:41:00.0                  # one or more PCI addresses providing this type

kolla-ansible -i ./multinode reconfigure --tags nova

# Flavor: request one virtual GPU
openstack flavor create --vcpus 8 --ram 32768 --disk 100 g1.vgpu.1
openstack flavor set g1.vgpu.1 --property "resources:VGPU"="1" --property "aggregate_instance_extra_specs:gpu"="true"

Nova creates the mediated device on demand when the instance boots. Install the exact guest driver validated against the vGPU Manager, guest OS, CUDA stack and selected vGPU release. The client token configures access to the NVIDIA licence service; monitor licence acquisition and test loss of connectivity rather than assuming one universal degradation mode.

6. Verify with Placement

VGPU inventory appears on a child resource provider of the compute node. For passthrough devices, report_in_placement publishes PCI inventory as CUSTOM_PCI_VENDOR_PRODUCT resources; the scheduler uses Placement allocations only when pci_in_placement is also enabled consistently. PciPassthroughFilter remains required.

# Resource providers for a GPU node and their inventories
openstack resource provider list --name gpu-node-01
openstack resource provider list --in-tree <GPU_NODE_RP_UUID>          # child providers (one per mdev-capable PCI device)
openstack resource provider inventory list <CHILD_RP_UUID>             # resource_class VGPU or CUSTOM_PCI_10DE_2331, total / used

# What the scheduler would find for a flavor
openstack allocation candidate list --resource VGPU=1
openstack allocation candidate list --resource CUSTOM_PCI_10DE_2331=1

# After boot: the instance's allocations
openstack resource provider allocation show <SERVER_UUID>

7. Pitfalls to test on the pilot

  • ■NUMA affinity: by default Nova requires the GPU and the instance CPUs on the same NUMA node. On dense hosts this blocks scheduling; numa_policy preferred trades a little latency for placeability. Measure both.
  • ■No live migration for passthrough instances (and constrained for vGPU): plan host maintenance as stop / start windows and tell the tenants.
  • ■IOMMU groups: if the GPU shares a group with another device (audio function, a NIC on the same PCIe switch), all of them must be bound to vfio-pci or the passthrough fails. Check before ordering servers.
  • ■Driver and firmware matrix: host vGPU driver, guest driver, CUDA and framework versions must match the vendor's compatibility table; pin them in your images and upgrade them together.
  • ■Licensing: vGPU without a reachable licence server degrades silently. Monitor the licence state from inside the guests and alert on it.
  • ■Observability: export GPU metrics (DCGM exporter in the guests or on bare metal) into the same Prometheus chain as the platform; a GPU at 0 % on a busy node is a scheduling bug, not idle capacity.
  • ■Next layer: once instances boot with a working GPU, provision Kubernetes on GPU flavors and install the NVIDIA GPU Operator; serve models with vLLM. That is a separate guide once we have run it.
Read the observability guide →