docs(monitors): heartbeat monitors and metric alerts guides; mask ping tokens in bundled nginx log
Deploy / deploy (push) Successful in 4m3s

This commit is contained in:
2026-09-17 09:46:49 +00:00
parent f230b66384
commit 53629450e0
7 changed files with 279 additions and 4 deletions
+111
View File
@@ -0,0 +1,111 @@
---
id: metric-alerts
title: Metric alerts
sidebar_label: Metric alerts
---
Agents already report disk, memory, load, reboot status and the state of every
container and systemd unit. A metric alert turns those reports into incidents:
"disk above 90% on any production server", "a systemd unit failed", "a reboot
has been pending for more than 7 days".
One rule covers as many servers as its tags match. There is nothing to install:
alerts use what agents already send, on any agent version.
## Creating a rule
1. Go to **Monitors** and choose **New monitor**.
2. Pick **Server metric**.
3. Under **Servers**, choose which servers by tag. Leave it empty to watch every server.
4. Choose the **Metric** (the condition below) and its **Threshold**.
5. Set **For (minutes)**: how long the condition must hold before it alerts.
6. Attach [notification channels](./notification-channels.md) and save.
Tags are resolved each time the rule is checked, so a server tagged `env=prod`
tomorrow is covered from then on, and a server whose tag is removed stops being
watched.
## Conditions
| Metric | Alerts when | Threshold |
| --------------------- | ---------------------------------------------------------- | --------- |
| Disk used | Used space is at or above the threshold | percent |
| Disk free below | Free space is at or below the threshold | GB |
| Memory used | Used memory is at or above the threshold | percent |
| Load per core | 1-minute load divided by CPU cores is at or above it | ratio |
| Systemd unit failed | Any unit on the server is `failed` | none |
| Container unhealthy | Any container's health check reports `unhealthy` | none |
| Reboot pending for | The server has needed a reboot for at least this long | days |
| Agent offline for | The agent has not been seen for at least this long | minutes |
The two disk metrics take an optional **Mount**, such as `/var`. Without one,
every mount is checked and the fullest one is reported.
## How a rule decides
Rules are checked every 30 seconds, against each matching server separately.
- When a server first meets the condition, it becomes **pending**.
- If the condition is still met after **For (minutes)**, the server goes
**down**, an incident opens for that server, and the channels fire.
- As soon as the condition clears, the server is **up** and its incident
closes. A pending server that clears never alerts.
"For" is continuous: a server that drops below the threshold for one check
starts the count again. Use it to ignore a nightly backup filling a disk for
ten minutes, or a load spike during a deploy. Set it to 0 to alert on the first
check.
### Per-server incidents
Each server has its own state and its own incident. If a rule matches 40
servers and 3 are breaching, you get 3 incidents, and each alert names its
server:
```text
[Vantage] Disk full (metric) on web-01 is DOWN: /var 94.2% used
```
The monitors list shows the rule as a whole, for example
`3 of 40 servers breaching`. The rule's own page lists every matching server
with its state, the current value and how long it has been breaching.
Webhook payloads carry a `server_name` field for metric alerts; see
[notification channels](./notification-channels.md#webhook).
## Stale data
Agents send metrics every 30 seconds. If a server's last metrics are more than
5 minutes old, the rule skips that server and keeps its previous state. A
powered-off server does not suddenly clear, or trip, a disk alert.
To be told about the server itself going quiet, add an **Agent offline for**
rule. It is the one metric that does not need fresh reports.
Some reports arrive less often:
- **Reboot pending** is refreshed with the agent's full inventory, every 15
minutes and at agent start.
- **Units and containers** are reported every 60 seconds.
## When servers leave a rule
If a server stops matching (its tag changed, or it was deleted) while it has an
open incident, the incident is closed quietly, with no recovery message. Nothing
recovered; the server simply stopped being watched.
## Tag-restricted API keys
An [API key restricted to tags](../reference/api-tokens.md#tag-restrictions) can
only create a metric rule whose tags include all of its own. A key restricted
to `env=staging` can create a rule for `env=staging role=web`, but not one for
every server, and not one for `env=prod`.
The same key cannot rename, disable, change or delete an existing rule that
reaches further than its tags, and it only sees incidents and per-server states
for servers inside them.
## Status pages
A metric rule can be added to a [status page](./status-pages.md) like any other
monitor. Its uptime reflects whether any matching server was down.