C
CYSKA
Consulting
EN FR
Navigation
OpenStack · Deployment

Deploy and upgrade an HA OpenStack cluster with Kolla-Ansible

Complete workflow for a production cluster spread across two availability zones: multinode inventory, globals.yml, prechecks, initial deployment, post-deployment validation, then release-by-release upgrades. Commands are illustrative and must be adapted to your release, network layout and storage backend.

Audience: platform engineers Reference: 3 controllers · 2 AZ · 20 computes · Ceph · OVN Updated: September 2026

1. Multinode inventory and availability zones

Kolla-Ansible groups hosts by role. The control plane runs on three controllers; compute nodes are split into two groups that will later map to two Nova availability zones. Keeping AZ membership visible in the inventory makes the mapping to host aggregates explicit and reproducible.

./multinode (simplified excerpt)

[control]
control-01
control-02
control-03

[network:children]
control

[compute:children]
az1-computes
az2-computes

[az1-computes]
az1-node-[01:10]

[az2-computes]
az2-node-[01:10]

# Nova availability zones are mapped to host aggregates after deployment (see section 4)

2. Passwords and globals.yml

Generate the service passwords once and keep /etc/kolla/passwords.yml under strict access control. A first deployment must explicitly define at least the base distribution, internal VIP, management and external network interfaces, storage backends and Neutron plugin. Pin the Kolla-Ansible package to the target stable branch; override openstack_release only when using a deliberate custom image versioning policy.

# 2.1 Generate the cluster passwords
kolla-genpwd

# 2.2 Key settings in /etc/kolla/globals.yml
kolla_internal_vip_address: "10.0.0.254"
kolla_base_distro: "ubuntu"
network_interface: "bond0"
neutron_external_interface: "bond1"
enable_haproxy: "yes"
enable_cinder: "yes"
cinder_backend_ceph: "yes"
glance_backend_ceph: "yes"
nova_backend_ceph: "yes"
neutron_plugin_agent: "ovn"

# 2.3 Pre-deployment checks on the target hosts
kolla-ansible -i ./multinode prechecks

External Ceph is enabled independently for Cinder, Glance and Nova. Follow the Kolla-Ansible directory layout for each service under /etc/kolla/config: dedicated ceph.conf files and keyrings for glance, cinder-volume, cinder-backup and nova. The Ceph cluster itself is deployed and operated separately, typically with cephadm.

3. Bootstrap, prechecks and deploy

bootstrap-servers prepares the hosts (Docker or Podman, users, kernel settings). prechecks must pass on every host before deploy: a failed precheck is far cheaper than a half-deployed control plane.

# First deployment
kolla-ansible -i ./multinode bootstrap-servers
kolla-ansible -i ./multinode prechecks
kolla-ansible -i ./multinode deploy

4. Post-deployment validation

post-deploy writes the admin credentials. Check that every agent is up and enabled before declaring the platform ready, then create the host aggregates that turn the two compute groups into Nova availability zones.

# Generate and load the admin environment variables
kolla-ansible -i ./multinode post-deploy
source /etc/kolla/admin-openrc.sh

# Check OpenStack agent health
openstack compute service list
openstack network agent list
openstack volume service list

# Map compute groups to Nova availability zones
openstack aggregate create --zone AZ1 az1
openstack aggregate create --zone AZ2 az2
for i in $(seq -w 1 10); do openstack aggregate add host az1 az1-node-$i; done
for i in $(seq -w 1 10); do openstack aggregate add host az2 az2-node-$i; done
openstack availability zone list --compute

5. Supported in-place upgrade path

kolla-ansible upgrade rolls newer container images out on the existing cluster without building a parallel environment. Use adjacent releases, or a skip-level upgrade between two consecutive SLURP releases only when Kolla-Ansible and every service in the deployment support that path. A Rocky to 2024.1 jump is not a supported single-step upgrade. Each hop requires its release notes and service-specific preparations, including RabbitMQ preparation for a supported skip-level path.

# In-place upgrade through a path supported by the target release notes
cd /etc/kolla

# 0. Back up the control plane database before touching anything
kolla-ansible -i ./multinode mariadb_backup

# 1. Upgrade the kolla-ansible package to the next release and refresh its dependencies
pip install --upgrade 'git+https://opendev.org/openstack/kolla-ansible@stable/<NEXT_RELEASE>'
kolla-ansible install-deps

# 2. Update the target release in globals.yml
# openstack_release: "<NEXT_RELEASE>"

# 3. Pull the new container images, then re-run the prechecks
kolla-ansible -i ./multinode pull
kolla-ansible -i ./multinode prechecks

# 4. Roll out the upgrade across control plane and compute nodes
kolla-ansible -i ./multinode upgrade

# 5. Validate before moving to the next release
source /etc/kolla/admin-openrc.sh
openstack compute service list
openstack network agent list

Not the default path for large release gaps

An in-place upgrade changes the live control plane and compute services. A failed step can extend the maintenance window or disrupt APIs and workloads. For several releases of gap, a parallel platform with volume migration is usually safer: see the Cinder manage / unmanage guide. We only keep the in-place path after validating backups, rollback, application resilience, compatibility between each release and explicit risk acceptance.

Read the Cinder manage / unmanage guide →

6. Pitfalls to check first

  • ■Read the release notes of every intermediate release: deprecated options in globals.yml break prechecks, and some releases change the default container runtime or database version.
  • ■Pin the kolla-ansible package to the stable branch matching openstack_release; a mismatch produces images that do not exist or silently pulls a newer release.
  • ■Keep the internal and external VIPs, the Ceph keyrings and passwords.yml under version control with encryption (for example Ansible Vault); losing passwords.yml means losing the cluster.
  • ■Validate live migration between the two AZ groups after deployment: shared Ceph storage makes it possible, but the Nova scheduler still needs consistent CPU models across hosts.
  • ■Wire observability before opening the platform to tenants: exporters, alert rules and dashboards are part of the deployment, not a later phase.
Read the observability guide →