OpenStack observability: Prometheus, Alertmanager and Node Exporter tuning
Configuration examples for the monitoring chain we deploy on OpenStack and Ceph platforms: openstack-exporter scraped by Prometheus, dashboards provisioned in Grafana, alerts routed by Alertmanager, and a Node Exporter tuned so that virtual interfaces and container bridges stop polluting dashboards. The last section shows how to publish custom metrics from scripts with the textfile collector.
1. Scrape OpenStack and hosts with Prometheus
Two families of metrics matter: what the OpenStack APIs report (service state, quotas, instance counts, API latency) through openstack-exporter, and what each host reports through node_exporter. Label hosts by role at scrape time so that dashboards and alerts can distinguish a controller from a compute node or a Ceph OSD host.
# Install Node Exporter on every OpenStack host (Debian/Ubuntu package, runs as user "prometheus")
apt-get install prometheus-node-exporter -y
# prometheus.yml (excerpt)
scrape_configs:
# openstack-exporter (default port 9180) talks to the OpenStack APIs with a clouds.yaml
- job_name: 'openstack-api'
static_configs:
- targets: ['controller-01:9180']
# Hosts, labelled by role
- job_name: 'node'
static_configs:
- targets: ['control-01:9100', 'control-02:9100', 'control-03:9100']
labels: { role: 'controller' }
- targets: ['az1-node-01:9100', 'az1-node-02:9100']
labels: { role: 'compute', az: 'AZ1' }
- targets: ['ceph-osd-01:9100', 'ceph-osd-02:9100']
labels: { role: 'ceph' }
# Ceph MGR prometheus module
- job_name: 'ceph'
static_configs:
- targets: ['ceph-mon-01:9283']
2. Provision Grafana dashboards as code
Dashboards belong in version control, not in a click-built Grafana instance. A file provider loads JSON dashboards from a directory at start-up; the same directory is deployed by Ansible with the rest of the platform. Typical panels: API latency per service, instance spawn failures, hypervisor CPU and memory, Ceph health and PG states, per-role network throughput.
# /etc/grafana/provisioning/dashboards/openstack.yml
apiVersion: 1
providers:
- name: 'OpenStack'
folder: 'Cloud'
type: file
disableDeletion: true
updateIntervalSeconds: 60
options:
path: /etc/grafana/provisioning/dashboards/json
3. Alert rules and Alertmanager routing
An alert is only useful if someone owns it. Each rule carries a severity and a service label; Alertmanager routes on those labels to the team that can act, with grouping to avoid one notification per host during an incident.
# /etc/prometheus/rules/openstack.yml
groups:
- name: openstack-critical
rules:
- alert: OpenStackExporterDown
expr: up{job="openstack-api"} == 0
for: 5m
labels:
severity: critical
service: openstack
annotations:
summary: Prometheus cannot scrape openstack-exporter
- alert: NovaComputeServiceDown
expr: openstack_nova_agent_state{adminState="enabled"} == 0
for: 10m
labels:
severity: critical
service: nova
annotations:
summary: "nova-compute down on {{ $labels.hostname }}"
- alert: NovaAgentMetricsMissing
expr: absent(openstack_nova_agent_state)
for: 10m
labels:
severity: critical
service: nova
annotations:
summary: openstack-exporter no longer exposes Nova agent metrics
- alert: CephHealthWarning
expr: ceph_health_status{job="ceph"} != 0
for: 10m
labels:
severity: warning
service: ceph
- alert: HostRootFilesystemAlmostFull
expr: (node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) < 0.10
for: 15m
labels:
severity: warning
service: host
# /etc/alertmanager/alertmanager.yml (excerpt)
route:
receiver: 'ops-team'
group_by: ['alertname', 'service']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
routes:
- matchers: [ service="ceph" ]
receiver: 'storage-team'
continue: true
- matchers: [ severity="critical" ]
receiver: 'oncall'
receivers:
- name: 'ops-team'
slack_configs:
- api_url: 'https://hooks.slack.com/services/...'
channel: '#cloud-ops'
- name: 'oncall'
webhook_configs:
- url: 'https://incident-tool.example.org/webhook/prometheus'
- name: 'storage-team'
email_configs:
- to: 'storage@example.org'
4. Node Exporter: control network-interface cardinality
OpenStack hosts expose many interfaces, but tap*, qbr*, qvb*, qvo*, br-int and br-tun belong to the Neutron datapath. Excluding them reduces cardinality but also removes workload traffic and useful failure signals. Keep collection conservative, then filter routine dashboards with PromQL or Grafana variables validated for each node role.
# /etc/default/prometheus-node-exporter
ARGS='--web.listen-address=":9100" \
--collector.netdev.device-exclude="^(lo|ovs-system)$" \
--collector.netclass.ignored-devices="^(lo|ovs-system)$"'
systemctl restart prometheus-node-exporter
# Keep tap*, qbr*, qvb*, qvo*, qr-*, qg-*, ha-*, br-int, br-tun and br-ex:
# they carry instance, router, tunnel, HA or provider traffic in Neutron.
# Verify what is still exported
curl -s localhost:9100/metrics | grep '^node_network_receive_bytes_total' | cut -d'"' -f2 | sort -u
Operational recommendation
Start by excluding only lo and the ovs-system placeholder. Review the remaining series by node role, then move display-only filtering into dashboards. Roll any collector-side exclusion out on one host per role before the whole fleet.
5. Custom metrics with the textfile collector
Node Exporter cannot natively scrape everything: results from cronjobs, backup scripts, certificate expiry checks or custom health scripts are not exposed by default. The textfile collector lets any script write metrics to a .prom file, which Node Exporter reads and forwards into the Prometheus / Grafana chain like any other metric.
# /etc/default/prometheus-node-exporter
ARGS='--web.listen-address=":9100" \
--collector.textfile.directory="/var/lib/node_exporter/custom_metrics" \
--collector.netdev.device-exclude="^(lo|ovs-system)$"'
# The Debian/Ubuntu package runs as the "prometheus" user
mkdir -p /var/lib/node_exporter/custom_metrics
chown prometheus:prometheus /var/lib/node_exporter/custom_metrics
systemctl restart prometheus-node-exporter
# /usr/local/bin/backup-with-metric.sh
#!/bin/bash
DIR=/var/lib/node_exporter/custom_metrics
/usr/local/bin/backup.sh \
&& echo "backup_last_success_timestamp $(date +%s)" > "$DIR/backup.prom.$$" \
&& mv "$DIR/backup.prom.$$" "$DIR/backup.prom"
# /etc/cron.d/backup-status (one job per line, no line continuation)
0 3 * * * root /usr/local/bin/backup-with-metric.sh
# Matching alert: no successful backup for 36 hours
# - alert: BackupStale
# expr: (time() - backup_last_success_timestamp > 36 * 3600) or absent(backup_last_success_timestamp)
# for: 1h
# labels: { severity: warning, service: backup }
When to use it
- Tracking the success or failure and duration of cronjobs and backup scripts
- Exposing certificate expiry dates or licence and quota checks
- Publishing results from custom health-check or maintenance scripts
- Feeding any one-off or batch job result into Prometheus / Grafana alerting and dashboards
Important: write to a temporary file and then rename (mv) it atomically into the textfile directory, otherwise Node Exporter can read a half-written file.
6. Pitfalls to check first
- ■openstack-exporter needs a dedicated Keystone user with read-only roles and its own clouds.yaml; never scrape with the admin credentials used for operations.
- ■Alert on symptoms first (API latency, spawn failures, degraded PGs) and on causes second (CPU, disk); a saturated CPU with healthy APIs is a capacity ticket, not a page.
- ■Review the exclusion regex after every Neutron or Kolla change: a renamed bridge silently reappears in dashboards, or worse, a production bond gets filtered.
- ■Version alert rules and dashboards with the platform code and test rule syntax in CI (promtool check rules) before deploying.
- ■Schedule a tuning workshop after two to four weeks of real production traffic: thresholds chosen on day one are always wrong somewhere.