C
CYSKA
Consulting
EN FR
Navigation
Observability · Prometheus

OpenStack observability: Prometheus, Alertmanager and Node Exporter tuning

Configuration examples for the monitoring chain we deploy on OpenStack and Ceph platforms: openstack-exporter scraped by Prometheus, dashboards provisioned in Grafana, alerts routed by Alertmanager, and a Node Exporter tuned so that virtual interfaces and container bridges stop polluting dashboards. The last section shows how to publish custom metrics from scripts with the textfile collector.

Audience: SRE and cloud operators Stack: Prometheus, Grafana, Alertmanager, node_exporter, openstack-exporter Updated: September 2026

1. Scrape OpenStack and hosts with Prometheus

Two families of metrics matter: what the OpenStack APIs report (service state, quotas, instance counts, API latency) through openstack-exporter, and what each host reports through node_exporter. Label hosts by role at scrape time so that dashboards and alerts can distinguish a controller from a compute node or a Ceph OSD host.

# Install Node Exporter on every OpenStack host (Debian/Ubuntu package, runs as user "prometheus")
apt-get install prometheus-node-exporter -y

# prometheus.yml (excerpt)
scrape_configs:
  # openstack-exporter (default port 9180) talks to the OpenStack APIs with a clouds.yaml
  - job_name: 'openstack-api'
    static_configs:
      - targets: ['controller-01:9180']

  # Hosts, labelled by role
  - job_name: 'node'
    static_configs:
      - targets: ['control-01:9100', 'control-02:9100', 'control-03:9100']
        labels: { role: 'controller' }
      - targets: ['az1-node-01:9100', 'az1-node-02:9100']
        labels: { role: 'compute', az: 'AZ1' }
      - targets: ['ceph-osd-01:9100', 'ceph-osd-02:9100']
        labels: { role: 'ceph' }

  # Ceph MGR prometheus module
  - job_name: 'ceph'
    static_configs:
      - targets: ['ceph-mon-01:9283']

2. Provision Grafana dashboards as code

Dashboards belong in version control, not in a click-built Grafana instance. A file provider loads JSON dashboards from a directory at start-up; the same directory is deployed by Ansible with the rest of the platform. Typical panels: API latency per service, instance spawn failures, hypervisor CPU and memory, Ceph health and PG states, per-role network throughput.

# /etc/grafana/provisioning/dashboards/openstack.yml
apiVersion: 1
providers:
  - name: 'OpenStack'
    folder: 'Cloud'
    type: file
    disableDeletion: true
    updateIntervalSeconds: 60
    options:
      path: /etc/grafana/provisioning/dashboards/json

3. Alert rules and Alertmanager routing

An alert is only useful if someone owns it. Each rule carries a severity and a service label; Alertmanager routes on those labels to the team that can act, with grouping to avoid one notification per host during an incident.

# /etc/prometheus/rules/openstack.yml
groups:
- name: openstack-critical
  rules:
  - alert: OpenStackExporterDown
    expr: up{job="openstack-api"} == 0
    for: 5m
    labels:
      severity: critical
      service: openstack
    annotations:
      summary: Prometheus cannot scrape openstack-exporter

  - alert: NovaComputeServiceDown
    expr: openstack_nova_agent_state{adminState="enabled"} == 0
    for: 10m
    labels:
      severity: critical
      service: nova
    annotations:
      summary: "nova-compute down on {{ $labels.hostname }}"

  - alert: NovaAgentMetricsMissing
    expr: absent(openstack_nova_agent_state)
    for: 10m
    labels:
      severity: critical
      service: nova
    annotations:
      summary: openstack-exporter no longer exposes Nova agent metrics

  - alert: CephHealthWarning
    expr: ceph_health_status{job="ceph"} != 0
    for: 10m
    labels:
      severity: warning
      service: ceph

  - alert: HostRootFilesystemAlmostFull
    expr: (node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) < 0.10
    for: 15m
    labels:
      severity: warning
      service: host
# /etc/alertmanager/alertmanager.yml (excerpt)
route:
  receiver: 'ops-team'
  group_by: ['alertname', 'service']
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h
  routes:
    - matchers: [ service="ceph" ]
      receiver: 'storage-team'
      continue: true
    - matchers: [ severity="critical" ]
      receiver: 'oncall'

receivers:
  - name: 'ops-team'
    slack_configs:
      - api_url: 'https://hooks.slack.com/services/...'
        channel: '#cloud-ops'
  - name: 'oncall'
    webhook_configs:
      - url: 'https://incident-tool.example.org/webhook/prometheus'
  - name: 'storage-team'
    email_configs:
      - to: 'storage@example.org'

4. Node Exporter: control network-interface cardinality

OpenStack hosts expose many interfaces, but tap*, qbr*, qvb*, qvo*, br-int and br-tun belong to the Neutron datapath. Excluding them reduces cardinality but also removes workload traffic and useful failure signals. Keep collection conservative, then filter routine dashboards with PromQL or Grafana variables validated for each node role.

# /etc/default/prometheus-node-exporter
ARGS='--web.listen-address=":9100" \
  --collector.netdev.device-exclude="^(lo|ovs-system)$" \
  --collector.netclass.ignored-devices="^(lo|ovs-system)$"'

systemctl restart prometheus-node-exporter

# Keep tap*, qbr*, qvb*, qvo*, qr-*, qg-*, ha-*, br-int, br-tun and br-ex:
# they carry instance, router, tunnel, HA or provider traffic in Neutron.

# Verify what is still exported
curl -s localhost:9100/metrics | grep '^node_network_receive_bytes_total' | cut -d'"' -f2 | sort -u

Operational recommendation

Start by excluding only lo and the ovs-system placeholder. Review the remaining series by node role, then move display-only filtering into dashboards. Roll any collector-side exclusion out on one host per role before the whole fleet.

5. Custom metrics with the textfile collector

Node Exporter cannot natively scrape everything: results from cronjobs, backup scripts, certificate expiry checks or custom health scripts are not exposed by default. The textfile collector lets any script write metrics to a .prom file, which Node Exporter reads and forwards into the Prometheus / Grafana chain like any other metric.

# /etc/default/prometheus-node-exporter
ARGS='--web.listen-address=":9100" \
  --collector.textfile.directory="/var/lib/node_exporter/custom_metrics" \
  --collector.netdev.device-exclude="^(lo|ovs-system)$"'

# The Debian/Ubuntu package runs as the "prometheus" user
mkdir -p /var/lib/node_exporter/custom_metrics
chown prometheus:prometheus /var/lib/node_exporter/custom_metrics

systemctl restart prometheus-node-exporter
# /usr/local/bin/backup-with-metric.sh
#!/bin/bash
DIR=/var/lib/node_exporter/custom_metrics
/usr/local/bin/backup.sh \
  && echo "backup_last_success_timestamp $(date +%s)" > "$DIR/backup.prom.$$" \
  && mv "$DIR/backup.prom.$$" "$DIR/backup.prom"

# /etc/cron.d/backup-status (one job per line, no line continuation)
0 3 * * * root /usr/local/bin/backup-with-metric.sh

# Matching alert: no successful backup for 36 hours
# - alert: BackupStale
#   expr: (time() - backup_last_success_timestamp > 36 * 3600) or absent(backup_last_success_timestamp)
#   for: 1h
#   labels: { severity: warning, service: backup }

When to use it

  • Tracking the success or failure and duration of cronjobs and backup scripts
  • Exposing certificate expiry dates or licence and quota checks
  • Publishing results from custom health-check or maintenance scripts
  • Feeding any one-off or batch job result into Prometheus / Grafana alerting and dashboards

Important: write to a temporary file and then rename (mv) it atomically into the textfile directory, otherwise Node Exporter can read a half-written file.

6. Pitfalls to check first

  • ■openstack-exporter needs a dedicated Keystone user with read-only roles and its own clouds.yaml; never scrape with the admin credentials used for operations.
  • ■Alert on symptoms first (API latency, spawn failures, degraded PGs) and on causes second (CPU, disk); a saturated CPU with healthy APIs is a capacity ticket, not a page.
  • ■Review the exclusion regex after every Neutron or Kolla change: a renamed bridge silently reappears in dashboards, or worse, a production bond gets filtered.
  • ■Version alert rules and dashboards with the platform code and test rule syntax in CI (promtool check rules) before deploying.
  • ■Schedule a tuning workshop after two to four weeks of real production traffic: thresholds chosen on day one are always wrong somewhere.