4.8 KiB
id, title, sidebar_label
| id | title | sidebar_label |
|---|---|---|
| metric-alerts | Metric alerts | Metric alerts |
Agents already report disk, memory, load, reboot status and the state of every container and systemd unit. A metric alert turns those reports into incidents: "disk above 90% on any production server", "a systemd unit failed", "a reboot has been pending for more than 7 days".
One rule covers as many servers as its tags match. There is nothing to install: alerts use what agents already send, on any agent version.
Creating a rule
- Go to Monitors and choose New monitor.
- Pick Server metric.
- Under Servers, choose which servers by tag. Leave it empty to watch every server.
- Choose the Metric (the condition below) and its Threshold.
- Set For (minutes): how long the condition must hold before it alerts.
- Attach notification channels and save.
Tags are resolved each time the rule is checked, so a server tagged env=prod
tomorrow is covered from then on, and a server whose tag is removed stops being
watched.
Conditions
| Metric | Alerts when | Threshold |
|---|---|---|
| Disk used | Used space is at or above the threshold | percent |
| Disk free below | Free space is at or below the threshold | GB |
| Memory used | Used memory is at or above the threshold | percent |
| Load per core | 1-minute load divided by CPU cores is at or above it | ratio |
| Systemd unit failed | Any unit on the server is failed |
none |
| Container unhealthy | Any container's health check reports unhealthy |
none |
| Reboot pending for | The server has needed a reboot for at least this long | days |
| Agent offline for | The agent has not been seen for at least this long | minutes |
The two disk metrics take an optional Mount, such as /var. Without one,
every mount is checked and the fullest one is reported.
How a rule decides
Rules are checked every 30 seconds, against each matching server separately.
- When a server first meets the condition, it becomes pending.
- If the condition is still met after For (minutes), the server goes down, an incident opens for that server, and the channels fire.
- As soon as the condition clears, the server is up and its incident closes. A pending server that clears never alerts.
"For" is continuous: a server that drops below the threshold for one check starts the count again. Use it to ignore a nightly backup filling a disk for ten minutes, or a load spike during a deploy. Set it to 0 to alert on the first check.
Per-server incidents
Each server has its own state and its own incident. If a rule matches 40 servers and 3 are breaching, you get 3 incidents, and each alert names its server:
[Vantage] Disk full (metric) on web-01 is DOWN: /var 94.2% used
The monitors list shows the rule as a whole, for example
3 of 40 servers breaching. The rule's own page lists every matching server
with its state, the current value and how long it has been breaching.
Webhook payloads carry a server_name field for metric alerts; see
notification channels.
Stale data
Agents send metrics every 30 seconds. If a server's last metrics are more than 5 minutes old, the rule skips that server and keeps its previous state. A powered-off server does not suddenly clear, or trip, a disk alert.
To be told about the server itself going quiet, add an Agent offline for rule. It is the one metric that does not need fresh reports.
Some reports arrive less often:
- Reboot pending is refreshed with the agent's full inventory, every 15 minutes and at agent start.
- Units and containers are reported every 60 seconds.
When servers leave a rule
If a server stops matching (its tag changed, or it was deleted) while it has an open incident, the incident is closed quietly, with no recovery message. Nothing recovered; the server simply stopped being watched.
Tag-restricted API keys
An API key restricted to tags can
only create a metric rule whose tags include all of its own. A key restricted
to env=staging can create a rule for env=staging role=web, but not one for
every server, and not one for env=prod.
The same key cannot rename, disable, change or delete an existing rule that reaches further than its tags, and it only sees incidents and per-server states for servers inside them.
Status pages
A metric rule can be added to a status page like any other monitor. Its uptime reflects whether any matching server was down.