docs: Document backup and restore

The page leads with KEY_ENCRYPTION_KEY rather than mentioning it in a
note, because holding a good database dump and no key is the way this goes
wrong.
This commit is contained in:
2026-09-07 14:29:03 +00:00
parent f0f2600bed
commit 89cc524f3f
3 changed files with 262 additions and 3 deletions
+68 -2
View File
@@ -563,6 +563,65 @@ inserted in front. The same setting also decides the address recorded in
`UpdateAgentCmd` carries a target version and Gitea base URL; the agent downloads and replaces itself.
### Backup and restore
`vantagectl` is a standalone Go module (`vantagectl/`), not a subcommand of
`server`. It needs its own module rather than living inside `server`'s for the
same reason `admin` and `sitesvc` already do: `server` imports the rest of
`server`'s dependency graph, and `spf13/cobra` has no business in a process
that also terminates gRPC streams and serves the REST API. More to the point,
`vantagectl` has to run when the control plane **does not** — a backup or
restore against a database with no server container alive at all — so it
cannot be a mode of the binary whose crash is the reason you need it.
The actual logic lives in `shared/backup` (dump, restore, verify, manifest,
fingerprint), not in `vantagectl/internal/cmd`, which holds only argument
parsing and operator-facing output. That split is what lets `server` import
`shared/backup` later — a scheduled in-process backup, say — without a second
implementation to keep in sync. `shared/cryptobox` is the same move one layer
down: it is now the **single** AES-256-GCM implementation, and
`server/internal/services/crypto.go` delegates to it rather than keeping its
own copy that `shared/backup` would otherwise have had to duplicate to decrypt
a probe value during `verify`.
**The archive stores a SHA-256 fingerprint of `KEY_ENCRYPTION_KEY`, never the
key.** `backup` refuses to run without the key set in the environment unless
`--allow-no-key` is passed, because an archive with no fingerprint at all
cannot later tell a restore that the wrong key is in hand — it can only find
that out when the data comes back as noise. The fingerprint is what turns that
failure into a refusal at `restore` time instead.
**Collections are enumerated live**`shared/backup` lists what the database
actually holds rather than reading `services.ScopedCollections`, the opposite
choice from the one instance-deletion purge makes. Purge must never miss a
tenant-scoped collection, so it keeps one hand-maintained registry; a backup
must never miss **any** collection, tenant-scoped or not (`migrations`,
`vulndb_meta`), so a static list is the wrong shape twice over — once for the
collections it would still owe `instance_id` deletion but not a backup, and
once for the two singleton collections that carry neither `instance_id` nor a
release note.
**Restore refuses a non-empty target database and has no merge semantics.**
There is no code path that upserts an archive's documents over existing ones:
merging two control planes' data reconciles nothing about which SSH keys are
still valid or which users still exist, and an upsert would resurrect a
revoked key or a deleted member from the older side. `--force` drops each
collection in the archive first, and is gated behind a typed confirmation
(the target database's name, typed back) on a terminal, or `--confirm-db NAME`
matching the target exactly with none. Naming the target in the command itself
means a copied command carries its intended target with it and cannot destroy
a different one by accident.
**`vantagectl/Dockerfile`'s runtime stage is `scratch`, and needs the same
explicit `/tmp` as `server/Dockerfile`.** `restore` extracts an archive to a
temporary directory before verifying its checksums, and a scratch image has no
`/tmp` for `os.MkdirTemp` to find — the same failure mode `vulnsched` hits on
`server`, but here it would break every restore rather than only vulnerability
scanning.
**`shared/` now fans out to four Go images** in `server-deploy.yml`:
`server`, `sitesvc`, `admin` and `vantagectl` — see the CI section below.
### API tokens and OpenAPI
A token is `vt_` plus 32 random bytes hex, shown once at creation and stored
@@ -1221,7 +1280,7 @@ GOOS=linux GOARCH=amd64 go build \
### `server-deploy.yml` — triggered on every push to `main`
Builds and pushes seven images to the Gitea container registry: `server`, `web`, `site`, `sitesvc`, `admin`, `adminsite` and `docsite`.
Builds and pushes eight images to the Gitea container registry: `server`, `web`, `site`, `sitesvc`, `admin`, `adminsite`, `docsite` and `vantagectl`.
Note that despite the name, **this workflow does not deploy** — it only builds and pushes. There is no SSH step. Rolling images out is a separate manual step on the host:
@@ -1237,9 +1296,16 @@ cd /opt/vantage && docker compose -f docker-compose.yml -f docker-compose.site.y
| `server` | `server/`, `shared/`, `proto/`, `go.work` |
| `admin` | `admin/`, `shared/`, `go.work` |
| `sitesvc` | `sitesvc/`, `shared/`, `go.work` |
| `vantagectl` | `vantagectl/`, `shared/`, `go.work` |
| `web` · `site` · `adminsite` · `docsite` | their own directory only |
`shared/` fans out to all three Go images because each of their Dockerfiles copies `shared/` from a root context — **if a fourth service ever imports `shared/`, add it to that list or it will ship stale**. A change to the workflow file rebuilds everything, since a build arg is baked into the image. So does anything that leaves no trustworthy base commit: a manual `workflow_dispatch`, a new branch, or a force-push whose old head is gone.
`shared/` fans out to **four** Go images (`server`, `sitesvc`, `admin`,
`vantagectl`) because each of their Dockerfiles copies `shared/` from a root
context — **if a fifth service ever imports `shared/`, add it to that list or
it will ship stale**. A change to the workflow file rebuilds everything, since
a build arg is baked into the image. So does anything that leaves no
trustworthy base commit: a manual `workflow_dispatch`, a new branch, or a
force-push whose old head is gone.
The gap this leaves: **changing a repo variable pushes no commit, so nothing rebuilds.** After editing `ADMIN_API_URL`, `HQ_URL` or `ADMIN_ENV`, run the workflow manually — that is what `workflow_dispatch` is there for. Base images also stop being refreshed on a service nobody touches; a periodic manual run covers that.