Skip to content
DBDeependra Bhatta~/notes
Observability#grafana · #systemd · #prometheus · #alertmanager · #monitoring

Prometheus Alerting With Alertmanager, Slack and Grafana

Write Prometheus alert rules, route them through Alertmanager to a Slack channel, add recovery alerts, and build a node_exporter dashboard in Grafana.

· updated · 16 min read
ON THIS PAGE

Metrics only help during an incident if someone is told about the problem and can see it on a dashboard. This guide turns the part 1 lab, where Prometheus on 192.168.56.70 scrapes node_exporter, Docker engine metrics and cAdvisor on 192.168.56.71, into a working alerting pipeline: alert rules, Alertmanager routing to Slack, recovery alerts and Grafana dashboards. It also covers two problems from the lab, a config file in the wrong location and an empty Slack title.

Prerequisites

  • The Prometheus lab from part 1, with all targets UP.
  • A Slack workspace where you can create apps and incoming webhooks.
  • sudo access on the Prometheus VM.

How alerting works in Prometheus

Prometheus only evaluates alert conditions; it cannot send emails or chat messages. Delivery belongs to a separate service, Alertmanager:

  1. Alerting rules in Prometheus define the conditions. Each rule is a PromQL expression; when it returns results, the alert is active.
  2. Prometheus sends active alerts to Alertmanager.
  3. Alertmanager groups them, silences or inhibits some, and turns the rest into notifications for receivers such as email, Slack, webhooks or PagerDuty.

Alerting rules

A rule lives in a rule group inside a rules file:

YMLrules.yaml
groups:
  - name: node
    rules:
      - alert: NodeDown
        expr: up{job="node"} == 0
        for: 5m

for requires the expression to stay true for that duration before the alert fires. A single failed scrape (a network blip, a scrape timeout) should not page anyone, so most rules wait a few minutes. During that time the alert moves through three states, visible on the Alerts page of the Prometheus UI:

StateMeaning
InactiveThe expression returns nothing. No problem
PendingThe expression returns results, but not yet for the full for duration. A problem is suspected but not yet confirmed
FiringTrue for longer than for. The alert is sent to Alertmanager

Labels classify an alert, and Alertmanager uses them to route and group:

YMLrules.yaml
groups:
  - name: node
    rules:
      - alert: NodeDown
        expr: up{job="node"} == 0
        labels:
          severity: warning
 
      - alert: MultipleNodesDown
        expr: avg without(instance) (up{job="node"}) <= 0.5
        labels:
          severity: critical

Annotations carry extra text such as a description. They are not part of the alert's identity, so they cannot be used for routing. They are Go templates: {{ .Labels }} (or $labels) gives the alert labels, {{ .Labels.instance }} one label, and {{ .Value }} the value that triggered it:

YMLrules.yaml
groups:
  - name: node
    rules:
      - alert: NodeFilesystemLowSpace
        expr: |
          100 * node_filesystem_free_bytes{job="node"} /
          node_filesystem_size_bytes{job="node"} < 70
        annotations:
          description: "Filesystem {{ .Labels.device }} on {{ .Labels.instance }} is low on space. Current available space is {{ .Value }}"

Install Alertmanager

Install Alertmanager 0.28.1 on the Prometheus VM with a script and a systemd unit, following the same pattern as Prometheus in part 1:

SHalertmanager_install.sh
#!/bin/bash
sudo useradd --no-create-home --shell /bin/false alertmanager
sudo mkdir /etc/alertmanager
 
wget https://github.com/prometheus/alertmanager/releases/download/v0.28.1/alertmanager-0.28.1.linux-amd64.tar.gz
tar xzf alertmanager-0.28.1.linux-amd64.tar.gz
cd alertmanager-0.28.1.linux-amd64
 
# Config goes to /etc/alertmanager, data to /var/lib/alertmanager
sudo mv alertmanager.yml /etc/alertmanager
sudo chown -R alertmanager:alertmanager /etc/alertmanager
sudo mkdir /var/lib/alertmanager
sudo chown -R alertmanager:alertmanager /var/lib/alertmanager
 
# Binaries: the server and amtool (the command-line client)
sudo cp alertmanager /usr/local/bin
sudo cp amtool /usr/local/bin
sudo chown alertmanager:alertmanager /usr/local/bin/alertmanager
sudo chown alertmanager:alertmanager /usr/local/bin/amtool
cd -
 
sudo cp alertmanager.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl start alertmanager
sudo systemctl enable alertmanager
sudo systemctl status alertmanager
INIalertmanager.service
[Unit]
Description=Alert Manager
Wants=network-online.target
After=network-online.target
 
[Service]
Type=simple
User=alertmanager
Group=alertmanager
ExecStart=/usr/local/bin/alertmanager \
       --config.file=/etc/alertmanager/alertmanager.yml \
       --storage.path=/var/lib/alertmanager
Restart=always
 
[Install]
WantedBy=multi-user.target

systemctl status alertmanager showing the Alert Manager service active (running)

The UI runs on port 9093. The Silences tab mutes alerts, as described below:

Alertmanager UI at 192.168.56.70:9093 on the Status page with cluster status ready and version info

Edit the config file the unit file points to, /etc/alertmanager/alertmanager.yml. In my lab, a copy created in /etc/prometheus by mistake caused a long debugging session later (see Troubleshooting).

ls in /etc/alertmanager showing alertmanager.yml in the correct directory

Write the alert rules

The rules watch the three jobs on the monitored VM from part 1. The job values must match the job_name entries in prometheus.yml exactly:

YMLrules.yaml
etcprometheusrules.yaml
groups:
  - name: my-alerts
    interval: 15s # how often this group is evaluated
    rules:
      - alert: NodeDown
        expr: up{job="Node_exporter_next_vm"} == 0
        for: 2m # wait 2 minutes before firing
        labels:
          team: infra
          env: dev
        annotations:
          message: "instance {{ .Labels.instance }} is currently down"
 
      - alert: Docker-engine
        expr: up{job="Docker_engine"} == 0
        for: 0m # fire immediately
        labels:
          team: microservices
          env: prod
        annotations:
          message: "{{ .Labels.instance }} is currently down"
 
      - alert: Frontend_App_Down
        expr: up{job="cAdvisor"} == 0
        for: 0m
        labels:
          team: Frontend
          env: dev
        annotations:
          message: "instance {{ .Labels.instance }} is currently down"

Next, load the rules and point Prometheus at Alertmanager. A relative path in rule_files resolves from the directory of prometheus.yml, so rules.yaml must sit in /etc/prometheus:

YMLprometheus.yml
global:
  scrape_interval: 15s
  evaluation_interval: 15s
 
alerting:
  alertmanagers:
    - static_configs:
        - targets:
            - 192.168.56.70:9093 # Alertmanager host and port
 
rule_files:
  - rules.yaml
 
scrape_configs:
  # same jobs as in part 1: prometheus, PrometheusVM,
  # Node_exporter_next_vm, Docker_engine, cAdvisor

Validate both files and restart Prometheus:

terminal
$ promtool check rules /etc/prometheus/rules.yaml
$ promtool check config /etc/prometheus/prometheus.yml
$ sudo systemctl restart prometheus

All three rules show up as inactive:

Prometheus Alerts page with the my-alerts group from /etc/prometheus/rules.yaml, three rules INACTIVE

To test the rules, stop Docker on the monitored VM. This also stops the cAdvisor container, so two targets go down:

terminal
$ sudo systemctl stop docker
Stopping 'docker.service', but its triggering units are still active:
docker.socket

Terminal running sudo systemctl stop docker with the docker.socket warning

The warning means docker.socket can start Docker again on the next API call; stop docker.socket as well to keep Docker down. Both alerts fire immediately because their for is 0m:

Prometheus Alerts page with Docker-engine and Frontend_App_Down FIRING and NodeDown inactive

Configure Alertmanager

alertmanager.yml has three main sections:

SectionPurpose
globalDefaults shared by all receivers (SMTP server, Slack API URL...). Can be overridden per receiver
routeA tree of rules that decides which receiver gets which alert
receiversNamed destinations, each with one or more notifiers (email, Slack, webhook...)

The first version sends everything to the default web.hook receiver:

YMLalertmanager.yml
etcalertmanageralertmanager.yml
route:
  group_by: ['alertname']
  group_wait: 30s      # wait before the first notification for a new group
  group_interval: 5m   # wait before notifying about new alerts in the same group
  repeat_interval: 1h  # resend if the alert is still firing
  receiver: 'web.hook' # default receiver
 
receivers:
  - name: 'web.hook'
    webhook_configs:
      - url: 'http://127.0.0.1:5001/'

With Docker still stopped, both alerts arrive in the Alertmanager UI, grouped by alertname and carrying the labels from the rules:

Alertmanager UI listing Docker-engine and Frontend_App_Down alerts under web.hook with env, job and team labels

Silences mute notifications for a time window, for example during a planned one-hour upgrade. Click Silence on an alert, or create one with matchers such as env="production":

Alertmanager New Silence form with start, 2h duration, end, matchers, creator and comment fields

Routing rules

  • Every config has one root route. Alerts that match no sub-route go to its receiver.
  • Sub-routes under routes: filter on labels. Alerts are checked against them in order.
  • Matching stops at the first sub-route that matches, unless that route has continue: true.
  • group_by takes label names. Alerts with the same values for those labels are sent as one notification. Without it, all alerts of a route are combined into a single notification.

There are two ways to match labels. match_re (and match) is the old form and is deprecated. matchers is the current form and supports =, !=, =~ (regex) and !~:

YMLalertmanager.yml
route:
  receiver: staff
  group_by: ['alertname', 'job']
  routes:
    - matchers:
        - job =~ "node|windows"
      receiver: infra-email
    - matchers:
        - job = "kubernetes"
      receiver: k8s-slack

To deliver one alert to several receivers, use continue: true:

YMLalertmanager.yml
route:
  receiver: default-receiver
  group_by: ['alertname', 'job']
  routes:
    - matchers:
        - job = "kubernetes"
      receiver: k8s-email
      continue: true # keep checking the next routes after this match
    - matchers:
        - severity = "critical"
      receiver: critical-slack

Here an alert with job="kubernetes" and severity="critical" goes to both k8s-email and critical-slack. Alerts that match neither go to default-receiver.

Like Prometheus, Alertmanager reads its config only at start. After an edit, validate the file, then restart or reload with any one of the last three commands:

terminal
$ amtool check-config /etc/alertmanager/alertmanager.yml
$ sudo systemctl restart alertmanager
$ sudo killall -HUP alertmanager
$ curl -X POST http://localhost:9093/-/reload

The Alertmanager configuration reference lists every receiver and option.

Send alerts to Slack

Slack accepts messages through an incoming webhook, a secret URL that posts to one channel (Slack docs).

  1. Create a workspace and a channel. This lab uses #alert-manager-practice in the devops practice workspace.

    Slack workspace devops practice with the #alert-manager-practice channel open

  2. On the Slack webhooks page, click Create an app and choose From scratch.

    Slack Create an app dialog with From a manifest and From scratch options

  3. Name the app and pick the workspace.

    Slack Name app and choose workspace form with app name alertmanager practice and workspace devops practice

  4. Open Incoming Webhooks, switch it on and click Add New Webhook.

    Slack app settings with Incoming Webhooks turned on and the Add New Webhook button

  5. Pick the channel (type its name if it is not listed) and click Allow. Copy the webhook URL that Slack displays.

    Slack channel picker for the new webhook with #alert-manager-practice selected

Add a Slack receiver and a sub-route for the three jobs:

YMLalertmanager.yml
etcalertmanageralertmanager.yml
route: # root route
  group_by: ['alertname']
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 1h
  receiver: 'web.hook' # default receiver
  routes:
    - matchers:
        - job =~ "Node_exporter_next_vm|Docker_engine|cAdvisor"
      group_by: ['team', 'env'] # label names, not values
      receiver: slack
 
receivers:
  - name: 'web.hook'
    webhook_configs:
      - url: 'http://127.0.0.1:5001/'
 
  - name: slack
    slack_configs:
      - api_url: https://hooks.slack.com/services/XXXXXXXXX/XXXXXXXXX/XXXXXXXXXXXXXXXXXXXXXXXX
        channel: '#alert-manager-practice'
        title: '{{ .GroupLabels.team }} has alerts in env: {{ .GroupLabels.env }}'
        text: '{{ range .Alerts }}{{ .Annotations.message }}{{ "\n" }}{{ end }}' # one line per alert

The default file also contains an inhibit_rules block, in which a critical alert mutes matching warning alerts. These rules carry no severity label, so the block has no effect and is omitted.

After an Alertmanager restart, the notifications reach the channel:

Slack channel #alert-manager-practice with alerts: 192.168.56.71:9323 and instance 192.168.56.71:8080 are currently down

The title reads only "has alerts in env:" with no team or env. The first run used group_by: ['prod', 'dev']. Those are label values, and no label is named prod or dev, so .GroupLabels was empty. group_by requires label names, ['team', 'env'], as in the config above.

Recovery alerts

To announce when a target comes back, rename the down alerts to CamelCase and add three recovery rules:

YMLrules.yaml
etcprometheusrules.yaml
groups:
  - name: my-alerts
    interval: 15s
    rules:
      # DOWN alerts
      - alert: NodeDown
        expr: up{job="Node_exporter_next_vm"} == 0
        for: 2m
        labels:
          team: infra
          env: dev
        annotations:
          message: "🚨 instance {{ $labels.instance }} is currently DOWN"
 
      - alert: DockerEngineDown
        expr: up{job="Docker_engine"} == 0
        for: 0m
        labels:
          team: microservices
          env: prod
        annotations:
          message: "🚨 Docker engine on {{ $labels.instance }} is currently DOWN"
 
      - alert: FrontendAppDown
        expr: up{job="cAdvisor"} == 0
        for: 0m
        labels:
          team: Frontend
          env: dev
        annotations:
          message: "🚨 instance {{ $labels.instance }} (frontend) is currently DOWN"
 
      # RECOVERY alerts: fire only if the target is up now and changed state recently
      - alert: NodeRecovered
        expr: up{job="Node_exporter_next_vm"} == 1 and changes(up{job="Node_exporter_next_vm"}[5m]) > 0
        for: 0m
        labels:
          team: infra
          env: dev
        annotations:
          message: "✅ instance {{ $labels.instance }} has RECOVERED (UP)"
 
      - alert: DockerEngineRecovered
        expr: up{job="Docker_engine"} == 1 and changes(up{job="Docker_engine"}[5m]) > 0
        for: 0m
        labels:
          team: microservices
          env: prod
        annotations:
          message: "✅ Docker engine on {{ $labels.instance }} has RECOVERED (UP)"
 
      - alert: FrontendAppRecovered
        expr: up{job="cAdvisor"} == 1 and changes(up{job="cAdvisor"}[5m]) > 0
        for: 0m
        labels:
          team: Frontend
          env: dev
        annotations:
          message: "✅ instance {{ $labels.instance }} (frontend) has RECOVERED (UP)"

The recovery expression breaks down as follows:

PQLPromQL
up{job="Docker_engine"} == 1 and changes(up{job="Docker_engine"}[5m]) > 0
PartMeaning
up{job="Docker_engine"} == 1The target is up right now
up{job="Docker_engine"}[5m]Its up/down history for the last 5 minutes (a range vector)
changes(...) > 0The value changed at least once in that window
andBoth must be true

The rule therefore fires only for a recent recovery and goes quiet once the target has been up for 5 minutes. After Docker starts again, all three recovery alerts fire:

Prometheus Alerts page with NodeRecovered, DockerEngineRecovered and FrontendAppRecovered FIRING

Slack message listing Docker engine, node and frontend instances on 192.168.56.71 as RECOVERED (UP)

Explore metrics in the Prometheus UI

Before Grafana, the Prometheus UI is the quickest way to find metric names. The menu next to the query bar contains Explore metrics:

Prometheus query options menu with Explore metrics highlighted

Explore metrics list with ALERTS, builder_builds_failed_total, cadvisor_version_info and container_cpu metrics

Any metric can be graphed directly. This query plots container memory from cAdvisor:

PQLPromQL
container_memory_rss

Graph tab for container_memory_rss over 1 hour rising from about 16 MB to 20 MB

These graphs suit a quick check, not a dashboard; Grafana fills that role.

Install Grafana

Grafana runs as a Docker image, on Kubernetes, or as a package; see Set up Grafana for the options. On Ubuntu, use the APT repository, which also delivers updates, or a .deb from the download page. This lab uses the .deb:

terminal
$ sudo apt-get install -y adduser libfontconfig1 musl
$ wget https://dl.grafana.com/grafana-enterprise/release/12.1.1/grafana-enterprise_12.1.1_16903967602_linux_amd64.deb
$ sudo dpkg -i grafana-enterprise_12.1.1_16903967602_linux_amd64.deb

The package prints the commands to start it. The service is named grafana-server, not grafana:

terminal
$ sudo systemctl daemon-reload
$ sudo systemctl enable grafana-server
$ sudo systemctl start grafana-server
$ sudo systemctl status grafana
Unit grafana.service could not be found.
$ sudo systemctl status grafana-server

Enabling and starting grafana-server; status grafana fails with unit not found, status grafana-server shows running

Grafana listens on port 3000. The first login is admin / admin, and Grafana requires a new password immediately (sign-in docs):

Grafana login page with the default admin username and password filled in

Add Prometheus as a data source

In Connections → Data sources, add a Prometheus data source and set the server URL:

Grafana Prometheus data source with server URL http://192.168.56.70:9090/

The lab's Prometheus UI has no authentication, so no other fields are required. If your Prometheus uses basic auth or TLS, fill in those fields as well. Click Save & test.

Import a ready-made dashboard

Building panels by hand requires a PromQL query for each one. As a starting point, import a community dashboard from grafana.com/grafana/dashboards. For node_exporter, Node Exporter Full (ID 1860) is the standard choice. Dashboards can be imported by ID, by uploading the JSON, or by pasting the JSON; enter the ID:

Grafana Import dashboard page with ID 1860 entered and the Load button

Select the Prometheus data source and click Import:

Import dashboard step selecting the prometheus data source before clicking Import

Grafana can also manage users and its own alert rules.

Troubleshooting

Alerts do not reach Slack

In my lab, the Slack config was correct but no messages arrived. Two places show what Alertmanager is doing:

  • Status page in the Alertmanager UI. It prints the loaded config, which still showed the default config instead of the edited one:

    Alertmanager Status page Config section showing the loaded global settings

  • systemctl status alertmanager. It shows the --config.file path in use and the recent log lines:

    systemctl status alertmanager showing --config.file=/etc/alertmanager/alertmanager.yml and Notify attempt failed warnings

The service reads /etc/alertmanager/alertmanager.yml, but the edits were in a copy under /etc/prometheus. Moving the config to /etc/alertmanager/alertmanager.yml and restarting Alertmanager fixed delivery.

The Notify attempt failed warnings mean Alertmanager could not deliver to a receiver. A likely source is the default web.hook receiver, which posts to http://127.0.0.1:5001/, where nothing listens in this lab. Check each receiver URL, including the Slack webhook.

Other problems

SymptomCauseFix
Rules missing from the Alerts pagerules.yaml not next to prometheus.yml, or a typo in the file name (rules.yml vs rules.yaml)promtool check config reports missing rule files. Fix the path and restart
Alert never fires for a target that is downjob in expr does not match job_name exactlyCopy the job name from prometheus.yml or the Target health page
Alertmanager does not start after an editRoute points at an undefined receiver, or a YAML erroramtool check-config /etc/alertmanager/alertmanager.yml
Slack title shows empty valuesgroup_by lists label values (prod, dev)Use label names: group_by: ['team', 'env']
Docker comes back after systemctl stop dockerdocker.socket restarts it on demandAlso run sudo systemctl stop docker.socket
Unit grafana.service could not be foundThe service is named differentlysudo systemctl status grafana-server

Key takeaways

  • Prometheus decides when an alert fires; Alertmanager decides who hears about it and how.
  • Use for: to avoid alerts on a single failed scrape, and match job labels exactly to prometheus.yml.
  • Check the config Alertmanager loaded (Status page, --config.file in systemctl status) before debugging anything else.
  • group_by takes label names. Use matchers instead of the deprecated match_re, and continue: true to notify several receivers.
  • Treat a Slack webhook URL as a secret; api_url_file keeps it out of the config.
  • Grafana with a community dashboard such as ID 1860 provides a complete node_exporter view without writing PromQL by hand.

This is the final part of the series. Start from the beginning with Prometheus Setup: node_exporter, TLS, Docker and Relabeling.