docs: status pages

This commit is contained in:
2026-08-25 08:44:38 +00:00
parent fa67d839cd
commit 72c9492223
4 changed files with 201 additions and 1 deletions
+72 -1
View File
@@ -456,6 +456,76 @@ there are two copies — `agent/internal/grpc/pb` and `server/internal/grpc/pb`.
A message added to one must be added to the other and to the `.proto`, in the
same commit.
### Status pages
Two collections: `status_pages` is the page itself — title, banner, published
flag, and an ordered list of sections each holding entries that pair a
`monitor_id` with a per-page display name. `status_incidents` holds both
operator-authored incidents and maintenance windows, sharing one document
shape because they share a timeline, an impact and a set of affected
components; each carries an explicit `page_ids` rather than deriving it from
`affected_monitors`, because adding a monitor to a page later must not
retroactively republish that monitor's old incidents to a new audience.
**`services.assembleSnapshot` is the redaction boundary, and it is the only
one.** It takes a `snapshotInput` built from already-fetched
`models.Monitor`/`models.Rollup`/`models.Incident` documents and returns a
`StatusSnapshot` built entirely from a parallel, deliberately smaller
vocabulary (`PublicComponent`, `PublicIncident`, …) that has no field for a
target URL, host, port, expected status, keyword, failure message,
certificate expiry, latency, runner or notification channel — `models.Monitor`
itself never reaches an anonymous caller, only the handful of fields
`assembleSnapshot` chooses to copy out of it. Being a pure function of already-
fetched data (no DB calls inside it) is what makes the boundary testable
without a database, which is the only thing standing between an editor adding
a field to `PublicComponent` and that field being a hostname.
Monitor-detected outages are **derived at read time, never copied**: each
snapshot assembly reads recent `incidents` for the page's monitors and folds
them into the timeline alongside the authored ones. There is no second
incidents table for automatic ones and no reconciliation between two records
of the same outage. A maintenance window in progress **repaints how a day is
drawn, never the uptime number** — `buildDays` computes each day's up/down
state and the 90-day percentage from rollups first, and
`applyMaintenanceRepaint` only overwrites today's display state afterward, so
a component that stayed up throughout a maintenance window still shows as up
in its history.
The public route, `GET /public/status/:pageId`, is mounted on the gin **root**,
outside `/api`, on purpose: `/api` carries `auth.Middleware`, `RequireScopes`,
`RateLimitTokens` and `RequireActiveLicense` by virtue of where it is mounted,
and a public route living there would need four exemptions — each one a hole a
later change to any of those four could widen back open. A missing page, an
unpublished page, and a page on the wrong host all answer the same 404;
inventing a distinct code for "exists but unpublished" would itself leak that
the page exists. A lapsed licence or a tier lacking `status_pages` answers 200
with `available:false` and a `reason`, never a 403 or a blank page — the
reader is a member of the public who can do nothing about either condition and
deserves an explanation, not a browser error.
Assembled snapshots are cached in Redis for **30 seconds**, keyed per
instance and page, and every authoring write (`UpdateStatusPage`,
`DeleteStatusPage`, and every incident mutation) invalidates its page's entry
immediately rather than waiting out the TTL — an operator posting an update
mid-incident should not wonder for half a minute whether it saved. A cache
miss, on Redis being down or on any read error, degrades to reassembly rather
than an error: the status page has to survive the outage it exists to report.
The public endpoint itself is rate limited to **120 requests per minute per
client address**, answering 429 with `Retry-After`, on the same fixed-window
pattern as `RateLimitTokens`.
**`TRUSTED_PROXIES` is load-bearing for that limiter, not cosmetic.** `main.go`
always calls `gin.SetTrustedProxies` with it; left unset, gin trusts no proxy
and `c.ClientIP()` falls back to the direct peer address — which, sat behind a
real reverse proxy, is the proxy's own address for every visitor. The rate
limiter then keys on one address for the whole fleet of readers, and the first
burst of legitimate traffic during an incident is what trips it. Set it to the
proxy's real address or CIDR, not merely a private range guess; the shipped
compose file and Helm chart default it to the RFC1918 ranges, which is right
for their own bundled reverse proxy but wrong the moment another one is
inserted in front. The same setting also decides the address recorded in
`audit_logs` and `console_sessions`.
### Agent self-update
`UpdateAgentCmd` carries a target version and Gitea base URL; the agent downloads and replaces itself.
@@ -831,7 +901,7 @@ Paddle is merchant of record; `admin/internal/paddle` is a thin REST client (no
## MongoDB Collections
`servers` · `keys` · `assignments` · `orgs` · `users` · `auth_providers` · `settings` · `secrets` · `workflows` · `workflow_steps` · `workflow_runs` · `workflow_log_lines` · `workflow_log_seq` · `monitors` · `incidents` · `monitor_rollups` · `notification_channels` · `console_sessions` · `audit_logs` · `server_packages` · `vuln_findings` · `vuln_alert_rules` · `vulndb_meta` · `server_workloads` · `api_tokens` · `migrations`
`servers` · `keys` · `assignments` · `orgs` · `users` · `auth_providers` · `settings` · `secrets` · `workflows` · `workflow_steps` · `workflow_runs` · `workflow_log_lines` · `workflow_log_seq` · `monitors` · `incidents` · `monitor_rollups` · `notification_channels` · `console_sessions` · `audit_logs` · `server_packages` · `vuln_findings` · `vuln_alert_rules` · `vulndb_meta` · `server_workloads` · `api_tokens` · `status_pages` · `status_incidents` · `migrations`
Every document except `migrations` carries `org_id`. Struct definitions are the source of truth — see `server/internal/models/`.
@@ -940,6 +1010,7 @@ Windows: MSI built by CI (WiX), or `installer/setup.ps1` registering the agent a
| `PROXY_ADVERTISE_HOST` | no | default `server`; the hostname guacd resolves the control plane by, handed to guacd as the relay's address. Wrong here and every console session fails at connect |
| `PROXY_LISTEN_HOST` | no | default `0.0.0.0`; the interface the ephemeral relay listener binds |
| `APP_ROOT_LABEL` | no | default `vantage`; wrong value disables the host/session org guard |
| `TRUSTED_PROXIES` | no | comma-separated CIDRs or addresses gin trusts for `X-Forwarded-For`. Empty means trust none: `c.ClientIP()` falls back to the direct peer address, which behind a real reverse proxy is that proxy's own address for every visitor — the public status page's per-address rate limit then keys on one address for the whole fleet of readers. Also the address recorded in `audit_logs` and `console_sessions`. Compose and the Helm chart default it to the RFC1918 ranges, right for their own bundled proxy and wrong the moment another one is inserted in front |
| `POD_IP` | no | this pod's own address, set by the Helm chart from the downward API. **Takes precedence over `PROXY_ADVERTISE_HOST`** — a console relay listener belongs to one replica, and a Service address names all of them |
| `VANTAGE_MIGRATE_ONLY` | no | run schema setup (migrations, index builders, default-step seeding) and exit without serving. `GRPC_HOST` is not required in this mode. Set by the Helm chart's pre-upgrade Job |
| `VANTAGE_SKIP_MIGRATIONS` | no | serve without running schema setup, on the assumption a Job already did. Set by the chart's Deployment whenever `server.migrationJob.enabled`. Unset under Compose, where one process still migrates and then serves |