feat: Cleanup old docs
This commit is contained in:
@@ -1,307 +0,0 @@
|
||||
# Multiple auth providers
|
||||
|
||||
Date: 2026-08-03
|
||||
|
||||
## Problem
|
||||
|
||||
An instance can configure exactly one OIDC provider. `instance_oidc` holds one
|
||||
document per instance, `/auth/oidc/start` takes no argument, and `/login`
|
||||
renders an unconditional "Sign in with your instance's SSO" button whether or
|
||||
not anything is configured behind it. Customers who federate with more than one
|
||||
identity source cannot, and customers who federate with none are shown a button
|
||||
that leads to an error.
|
||||
|
||||
## Goals
|
||||
|
||||
- N auth providers per instance, each independently enabled and named.
|
||||
- Login page renders one button per enabled provider, and none when there are
|
||||
none.
|
||||
- Local email/password login can be turned off per instance.
|
||||
- Presets for the common identity providers, so a customer supplies a tenant ID
|
||||
rather than an issuer URL.
|
||||
- Existing configured SSO keeps working across the upgrade with no customer
|
||||
action.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- SAML. Different protocol, metadata parsing and certificate handling; not in
|
||||
this work.
|
||||
- Per-provider role or group mapping. Provisioned users remain `member`, as
|
||||
today.
|
||||
- Provider-specific account linking. An email address is an email address; the
|
||||
existing instance-scoped lookup stands.
|
||||
|
||||
## Data model
|
||||
|
||||
New collection `auth_providers`, one document per provider:
|
||||
|
||||
```go
|
||||
type AuthProvider struct {
|
||||
ID bson.ObjectID `bson:"_id,omitempty" json:"_id,omitempty"`
|
||||
InstanceID string `bson:"instance_id" json:"instance_id"`
|
||||
ProviderID string `bson:"provider_id" json:"provider_id"`
|
||||
Name string `bson:"name" json:"name"`
|
||||
Kind string `bson:"kind" json:"kind"` // "oidc" | "oauth2"
|
||||
Preset string `bson:"preset" json:"preset"` // "" for custom
|
||||
Issuer string `bson:"issuer" json:"issuer"`
|
||||
ClientID string `bson:"client_id" json:"client_id"`
|
||||
ClientSecretEnc string `bson:"client_secret_enc,omitempty" json:"-"`
|
||||
Scopes []string `bson:"scopes" json:"scopes"`
|
||||
Enabled bool `bson:"enabled" json:"enabled"`
|
||||
CallbackNotice bool `bson:"callback_notice" json:"callback_notice"`
|
||||
Order int `bson:"order" json:"order"`
|
||||
CreatedAt time.Time `bson:"created_at" json:"created_at"`
|
||||
UpdatedAt time.Time `bson:"updated_at" json:"updated_at"`
|
||||
}
|
||||
```
|
||||
|
||||
`ProviderID` is a short random identifier, not the Mongo `_id`: it appears in
|
||||
the callback URL a customer pastes into their IdP, and an `_id` there would
|
||||
publish a database key.
|
||||
|
||||
Unique index on `(instance_id, provider_id)`. Index build is fatal on failure,
|
||||
matching `EnsureAuthIndexes` — a duplicate `provider_id` within an instance
|
||||
would make the callback ambiguous.
|
||||
|
||||
`ClientSecretEnc` is AES-256-GCM under `KEY_ENCRYPTION_KEY`, as
|
||||
`instance_oidc.client_secret_enc` is today, and is never serialised.
|
||||
|
||||
### Presets
|
||||
|
||||
A Go table in `server/internal/auth/presets.go`, not database rows — adding one
|
||||
is a commit, not a migration.
|
||||
|
||||
| Preset | Kind | Issuer | Input asked of the customer | Default scopes |
|
||||
| ------- | -------- | ------------------------------------------------- | --------------------------- | ----------------------------- |
|
||||
| `entra` | `oidc` | `https://login.microsoftonline.com/{tenant}/v2.0` | Directory (tenant) ID | `openid profile email` |
|
||||
| `google`| `oidc` | `https://accounts.google.com` | none | `openid profile email` |
|
||||
| `okta` | `oidc` | `https://{domain}/oauth2/default` | Okta org domain | `openid profile email` |
|
||||
| `github`| `oauth2` | n/a | none | `read:user user:email` |
|
||||
| `` (custom) | `oidc` | supplied verbatim | Issuer URL | `openid profile email` |
|
||||
|
||||
The issuer template is expanded server-side on save; the stored `Issuer` is
|
||||
always the resolved URL, so nothing downstream has to know a preset existed.
|
||||
|
||||
### Settings
|
||||
|
||||
`settings.local_login_enabled bool`, defaulting true. Absent on existing
|
||||
documents, and Go's zero value for `bool` is false, so the field is read through
|
||||
a `*bool` and a nil pointer means enabled. A plain `bool` would silently
|
||||
disable password login on every instance in the fleet at upgrade.
|
||||
|
||||
## Migration
|
||||
|
||||
`0005_auth_providers` — the next free number; `0004_instance_rename` is the
|
||||
highest recorded today. For each document in `instance_oidc`, insert one
|
||||
`auth_providers` document:
|
||||
|
||||
- `Name: "Single sign-on"`
|
||||
- `Preset: ""`, `Kind: "oidc"`
|
||||
- `Issuer`, `ClientID`, `Enabled` copied
|
||||
- `ClientSecretEnc` copied **verbatim**, not decrypted and re-encrypted — a
|
||||
migration that needs `KEY_ENCRYPTION_KEY` fails on an instance that has none
|
||||
and strands the SSO configuration.
|
||||
- `Scopes: ["openid", "profile", "email"]`, matching what `oidc.go` hardcodes
|
||||
today.
|
||||
- `ProviderID` freshly generated.
|
||||
- `CallbackNotice: true` — this provider's redirect URI has changed and an
|
||||
administrator has not yet acknowledged it. Set only by the migration; cleared
|
||||
by the settings UI. New providers are created `false`.
|
||||
|
||||
`instance_oidc` is left in place and no longer read. Idempotent by skipping any
|
||||
instance that already has an `auth_providers` document, so a re-run after a
|
||||
partial failure completes rather than duplicating.
|
||||
|
||||
## Auth flow
|
||||
|
||||
Routes:
|
||||
|
||||
```
|
||||
GET /auth/oidc/:providerId/start
|
||||
GET /auth/oidc/:providerId/callback
|
||||
```
|
||||
|
||||
The old unparameterised `/auth/oidc/start` and `/auth/oidc/callback` are
|
||||
**removed**, not retained. See Upgrade impact below — this breaks configured SSO
|
||||
until the customer updates their IdP, and that is accepted deliberately rather
|
||||
than carried as a compatibility path.
|
||||
|
||||
The state token in Redis stores `{instance_id, provider_id}` rather than the
|
||||
bare instance ID. The callback resolves its provider from the consumed state
|
||||
and cross-checks it against `:providerId` in the path, refusing a mismatch —
|
||||
the path alone is attacker-controlled, and the state is the half that was
|
||||
issued by the start handler.
|
||||
|
||||
`providerForInstance` becomes `providerFor(ctx, c, instanceID, providerID)`.
|
||||
The `go-oidc` provider cache keys on `provider_id`, not instance. Saving,
|
||||
disabling or deleting a provider evicts that key.
|
||||
|
||||
`redirectURL(c, providerID)` returns the one per-provider shape, and returns the
|
||||
same URL in the start and callback halves of a flow — an IdP rejects the token
|
||||
exchange if they differ.
|
||||
|
||||
### OIDC providers
|
||||
|
||||
Unchanged from the current implementation: `AuthCodeURL` with the stored
|
||||
scopes, exchange, `id_token` verified against the provider's key set with
|
||||
`ClientID` as audience, `email` and `name` claims extracted.
|
||||
|
||||
### GitHub (`kind: "oauth2"`)
|
||||
|
||||
GitHub is OAuth2 and issues no `id_token`, so it takes a separate branch:
|
||||
exchange the code, then `GET https://api.github.com/user/emails` with the access
|
||||
token and take the address that is both `primary` and `verified`. An
|
||||
unverified-only response is refused — an unverified address is not proof of
|
||||
control, and accepting one would let anyone holding a GitHub account claim any
|
||||
address in the instance. `name` comes from `GET https://api.github.com/user`.
|
||||
|
||||
Both branches converge on one function:
|
||||
|
||||
```go
|
||||
func completeSSOLogin(c *gin.Context, instanceID, email, name string) error
|
||||
```
|
||||
|
||||
which holds today's lookup-or-provision, session creation, `TouchLastLogin` and
|
||||
cookie set, verbatim. Email is lower-cased before lookup, and the lookup stays
|
||||
`GetUserInInstanceByEmail` — instance-scoped, as it is now.
|
||||
|
||||
### Licence gate
|
||||
|
||||
`services.GetLicenseState(instanceID).Feature("oidc")` continues to gate both
|
||||
the start and the callback, for every provider kind, and is checked on the
|
||||
callback against the instance named by the consumed state rather than the host.
|
||||
Unchanged behaviour, applied to more providers.
|
||||
|
||||
## REST API
|
||||
|
||||
Unauthenticated:
|
||||
|
||||
```
|
||||
GET /auth/providers
|
||||
-> {"local_enabled": true,
|
||||
"providers": [{"id": "...", "name": "...", "preset": "entra"}]}
|
||||
```
|
||||
|
||||
Instance is resolved from the host, as `/auth/bootstrap-status` already does.
|
||||
The response carries **no issuer, no client ID and no secret** — it is served to
|
||||
anyone who can reach the login page.
|
||||
|
||||
Session-authed, `owner|admin`, under `/api`:
|
||||
|
||||
```
|
||||
GET,POST /auth/providers
|
||||
PUT,DELETE /auth/providers/:id
|
||||
POST /auth/providers/:id/test
|
||||
```
|
||||
|
||||
`test` fetches the provider's discovery document (or, for GitHub, calls the API
|
||||
with the stored credentials) and reports reachability. It does not sign anyone
|
||||
in.
|
||||
|
||||
`GET,PUT /api/org/oidc` is removed along with the old auth routes. Its only
|
||||
caller is `OIDCCard.tsx`, which this work replaces, and a compatibility shim
|
||||
over a one-of-many model would have to invent which provider it means.
|
||||
|
||||
Every mutation writes an audit event, as every mutating path does.
|
||||
|
||||
### Lockout guards
|
||||
|
||||
Both refused with 409 and a distinct error code:
|
||||
|
||||
- `local_login_required` — disabling local login while zero providers are
|
||||
enabled.
|
||||
- `last_provider` — disabling or deleting the last enabled provider while local
|
||||
login is off.
|
||||
|
||||
These are enforced in the service layer, not the handler, so the two endpoints
|
||||
that can reach the condition cannot disagree.
|
||||
|
||||
## Frontend
|
||||
|
||||
### Settings
|
||||
|
||||
`web/components/settings/OIDCCard.tsx` becomes `AuthProvidersCard`, in the
|
||||
Access group of `/settings` where the OIDC card already lives. It renders the
|
||||
provider list with per-row enable toggle, edit, delete and drag ordering, an
|
||||
Add flow that asks for the preset first and then only the fields that preset
|
||||
needs, and the local-login toggle beneath the list. A guard violation surfaces
|
||||
the 409's message rather than a generic failure.
|
||||
|
||||
Every provider row shows its **callback URL** with click-to-copy — that is the
|
||||
value the customer pastes into their IdP, it now differs per provider, and after
|
||||
the upgrade every migrated provider needs it re-pasted. A migrated provider
|
||||
additionally carries a warning until an administrator dismisses it, naming the
|
||||
change and the URL. Dismissal is per provider, stored on the document.
|
||||
|
||||
### Login page
|
||||
|
||||
`web/app/login/page.tsx` calls `/auth/providers` on mount alongside the existing
|
||||
`bootstrapStatus` call, and renders on the result:
|
||||
|
||||
| `local_enabled` | providers | Rendered |
|
||||
| --------------- | --------- | --------------------------------------------------- |
|
||||
| true | none | Password form only. No divider, no buttons. |
|
||||
| true | some | Password form, divider, one button per provider. |
|
||||
| false | some | Buttons only. No form, no divider. |
|
||||
| false | none | Password form (see below). |
|
||||
|
||||
The last row cannot be reached through the API — the guards above prevent it —
|
||||
but a hand-edited database could produce it, and a login page that renders
|
||||
nothing at all is unrecoverable without database access. It therefore falls back
|
||||
to the password form.
|
||||
|
||||
The current unconditional SSO button and its "SSO must be enabled for this
|
||||
instance by an administrator" note are both removed; the button now only exists
|
||||
when it works.
|
||||
|
||||
Buttons are labelled with the provider's `Name` and carry the preset's icon
|
||||
where there is one, a neutral key glyph otherwise. Presets never override the
|
||||
name — a customer who calls their Entra provider "Staff" gets "Staff".
|
||||
|
||||
Errors keep the existing `/login?error=<code>` redirect convention.
|
||||
|
||||
## Testing
|
||||
|
||||
- Migration: an `instance_oidc` document produces one enabled provider with the
|
||||
ciphertext byte-identical; a re-run inserts nothing further.
|
||||
- `local_login_enabled` absent decodes as enabled.
|
||||
- Guards: both 409 paths, and the enable/disable sequences that approach them
|
||||
without crossing.
|
||||
- Per-provider callback: two providers in one instance, each resolving to its
|
||||
own configuration; a `provider_id` from another instance answers 404.
|
||||
- A callback whose `:providerId` disagrees with the consumed state is refused,
|
||||
and the state is consumed rather than left replayable.
|
||||
- The removed routes (`/auth/oidc/start`, `/auth/oidc/callback`,
|
||||
`/api/org/oidc`) answer 404.
|
||||
- GitHub: primary+verified selected; verified-only-absent refused.
|
||||
- `/auth/providers` response contains no issuer, client ID or secret.
|
||||
|
||||
## Upgrade impact
|
||||
|
||||
**This release breaks configured SSO until each customer updates their identity
|
||||
provider.** The old `/auth/oidc/callback` is gone, migrated providers are
|
||||
reachable only at `/auth/oidc/<providerId>/callback`, and an IdP still pointing
|
||||
at the old URL fails the flow.
|
||||
|
||||
It is a deliberate trade: one callback shape rather than two, no
|
||||
`legacy_callback` branch through `redirectURL`, and no permanently retained
|
||||
route whose only purpose is a single past upgrade.
|
||||
|
||||
Mitigations, in order of who sees them first:
|
||||
|
||||
- The settings card shows the new callback URL per provider with click-to-copy,
|
||||
and a migrated provider carries a dismissable warning naming the change.
|
||||
- The failure is visible rather than silent: an IdP rejects the redirect URI
|
||||
before Vantage is reached, so the customer sees their own provider's error.
|
||||
- Local password login is unaffected, so no instance is locked out — an
|
||||
administrator can always sign in to fix the URL. This is why
|
||||
`local_login_enabled` defaults to true and why nothing in this migration
|
||||
turns it off.
|
||||
- Release notes and `docsite/docs/vantage/settings.md` state the required
|
||||
action.
|
||||
|
||||
## Deployment notes
|
||||
|
||||
No new environment variables. No agent change. `KEY_ENCRYPTION_KEY` is already
|
||||
required wherever OIDC was configured, and the migration does not add a
|
||||
dependency on it.
|
||||
@@ -1,235 +0,0 @@
|
||||
# Server tags and scheduled workflows
|
||||
|
||||
Date: 2026-08-04
|
||||
|
||||
Two features, designed together because the second is worth much less without
|
||||
the first. Tags make a target set describable; schedules make it recur. A
|
||||
nightly job that patches "everything tagged `env:staging`" needs both halves,
|
||||
and neither half is large on its own.
|
||||
|
||||
---
|
||||
|
||||
## Part A — Server tags
|
||||
|
||||
### Model
|
||||
|
||||
`models.Server` gains one field:
|
||||
|
||||
```go
|
||||
Tags map[string]string `bson:"tags,omitempty" json:"tags,omitempty"`
|
||||
```
|
||||
|
||||
Keys and values are lowercase `[a-z0-9_-]`. Keys are capped at 32 characters,
|
||||
values at 64, and a server holds at most 20 tags. Validation lives in the
|
||||
service layer rather than the handler, so the tag endpoint, the server-create
|
||||
path and anything added later cannot disagree about what a valid tag is.
|
||||
|
||||
There is **no `tags` collection.** A tag is a property of a server, not an
|
||||
entity with a lifecycle: a registry would need reference counting to know when
|
||||
a tag stopped existing, and garbage collection to act on it, which is work
|
||||
bought for nothing. The list of known keys and values that the UI offers for
|
||||
autocomplete is a distinct aggregation over `servers`, cached for 60 seconds —
|
||||
the same treatment org lookups already get.
|
||||
|
||||
No reserved keys ship in this change. If inventory-derived tags (`os`, `arch`)
|
||||
are added later they take a `sys:` key prefix, so a user tag written today can
|
||||
never collide with a system tag invented tomorrow.
|
||||
|
||||
Index: `{instance_id: 1, "tags.$**": 1}` — a wildcard index over the tag
|
||||
subdocument, because the queried key is chosen by the user at request time and
|
||||
cannot be named in advance.
|
||||
|
||||
### API
|
||||
|
||||
```
|
||||
PUT /api/servers/:id/tags # replace the whole map
|
||||
GET /api/servers/tags # known keys and values, for pickers
|
||||
GET /api/servers?tag=env:prod # repeatable; AND across keys
|
||||
```
|
||||
|
||||
`PUT` replaces the entire map rather than patching one tag. A tag set is small
|
||||
enough that sending all of it is free, and last-write-wins over a whole map is
|
||||
easier to reason about than merge semantics between two people editing the same
|
||||
server. The audit event records the map before and after.
|
||||
|
||||
`?tag=` is repeatable and ANDs: `?tag=env:prod&tag=role:web` matches servers
|
||||
carrying both. A malformed value (no colon, unknown characters) is a 400 rather
|
||||
than a silent empty result — a filter that matches nothing and a filter that is
|
||||
nonsense look identical in a list, and only one of them is the user's fault.
|
||||
|
||||
### Targeting
|
||||
|
||||
`models.Workflow` gains `TargetTags map[string]string` beside the existing
|
||||
`TargetServerIDs`. One function in `services` resolves them:
|
||||
|
||||
```go
|
||||
ResolveTargets(ctx, instanceID string, ids []string, tags map[string]string) ([]Server, error)
|
||||
```
|
||||
|
||||
- Result is the **distinct union** of the explicit IDs and the tag matches.
|
||||
- Tag matching ANDs across keys.
|
||||
- Offline servers are included. The dispatcher already answers 503 per server,
|
||||
and a patch run that silently omits an unreachable machine is worse than one
|
||||
that visibly fails on it.
|
||||
- Empty IDs **and** empty tags returns `ErrNoTargets` (400). A workflow that
|
||||
matches nothing must say so rather than report success over zero servers.
|
||||
|
||||
The resolved set is snapshotted into `WorkflowRun.ServerRuns` exactly as today.
|
||||
History records what actually ran, not what the selector would match when the
|
||||
run is later read back — the same reason `steps_snapshot` exists.
|
||||
|
||||
### Frontend
|
||||
|
||||
- **Server detail**: tag chips in the header with an inline editor. Keys
|
||||
autocomplete from `GET /api/servers/tags`, values autocomplete per key.
|
||||
- **`/servers`**: a filter bar that reads and writes the same `?tag=` query
|
||||
params the API takes, so a filtered fleet view is a URL someone can send.
|
||||
- **Workflow designer**: a target section holding both inputs, with a live
|
||||
"runs on 14 servers" readout that lists them on hover. The union model costs
|
||||
us the at-a-glance answer to "what will this touch"; this readout buys it
|
||||
back, and it is the reason the union is acceptable.
|
||||
|
||||
---
|
||||
|
||||
## Part B — Scheduled workflows
|
||||
|
||||
### Model
|
||||
|
||||
```go
|
||||
type Schedule struct {
|
||||
Enabled bool `bson:"enabled" json:"enabled"`
|
||||
Cron string `bson:"cron" json:"cron"` // 5-field
|
||||
TZ string `bson:"tz" json:"tz"` // IANA name
|
||||
}
|
||||
|
||||
type Skip struct {
|
||||
Reason string `bson:"reason" json:"reason"` // "missed" | "already_running"
|
||||
Due time.Time `bson:"due" json:"due"`
|
||||
At time.Time `bson:"at" json:"at"`
|
||||
}
|
||||
```
|
||||
|
||||
On `Workflow`:
|
||||
|
||||
```go
|
||||
Schedule *Schedule `bson:"schedule,omitempty"`
|
||||
NextRunAt *time.Time `bson:"next_run_at,omitempty"` // UTC, indexed
|
||||
LastRunAt *time.Time `bson:"last_run_at,omitempty"`
|
||||
LastSkipped *Skip `bson:"last_skipped,omitempty"`
|
||||
```
|
||||
|
||||
`next_run_at` is **persisted, not held in memory.** A leader handover between
|
||||
computing the next occurrence and firing it would otherwise either lose the
|
||||
occurrence or fire it twice. Coordination state has to live where every replica
|
||||
can see it — the same argument that put `workflow_log_seq` in MongoDB.
|
||||
|
||||
Cron parsing uses `robfig/cron/v3`'s **parser only** — `Parse` and
|
||||
`Next(time)`. Its scheduler and goroutines are not used; the loop below is ours
|
||||
and has to be, because it runs under the leader lock.
|
||||
|
||||
**Alpine ships no tzdata.** `server/Dockerfile` builds a slim image, so
|
||||
`time.LoadLocation("Europe/London")` returns an error and every schedule
|
||||
falls back to UTC — an hour wrong for half the year, in the direction nobody
|
||||
notices until a maintenance window lands in business hours. `main` therefore
|
||||
imports `_ "time/tzdata"`, embedding the database in the binary. Zone names are
|
||||
also validated at save time, so an unknown zone is a 400 rather than a surprise
|
||||
at 2am.
|
||||
|
||||
### Scheduler
|
||||
|
||||
A new `server/internal/workflowsched` package, started inside the **existing**
|
||||
`bus.RunAsLeader("housekeeping", …)` alongside `monitorsched`, `StartReaper`
|
||||
and the sweepers. One role, one lock. It takes the same cancellable context and
|
||||
returns the instant leadership is lost.
|
||||
|
||||
The loop ticks every 30 seconds:
|
||||
|
||||
1. `find({schedule.enabled: true, next_run_at: {$lte: now}})`.
|
||||
2. **Claim atomically.** `findOneAndUpdate` matching the document *and* its
|
||||
current `next_run_at`, setting the recomputed next occurrence. A process
|
||||
that reaches the same document after another has claimed it matches nothing
|
||||
and does nothing. The claim is what makes this correct; the leader lock only
|
||||
makes it cheap.
|
||||
3. **Grace check.** If `now - due > 1h`, record
|
||||
`last_skipped{reason: "missed"}`, write an audit event, and do not run. A
|
||||
job missed by ten minutes during a deploy should still run; one missed by
|
||||
two days should not fire at lunchtime.
|
||||
4. **Overlap check.** If a run for this workflow is still active, record
|
||||
`last_skipped{reason: "already_running"}`, audit, and do not run. A patch
|
||||
workflow must never run twice at once, and a silent skip is how a week goes
|
||||
by before anyone notices nothing ran.
|
||||
5. Otherwise start the run through the **same** `RunWorkflow` path a person
|
||||
uses, with `TriggeredBy: "schedule"`.
|
||||
|
||||
Step 5 is the design. A scheduled run is an ordinary run with a different
|
||||
trigger: no second dispatch path, no second snapshot format, and the run detail
|
||||
page needs no changes to display one.
|
||||
|
||||
### API
|
||||
|
||||
```
|
||||
PUT /api/workflows/:id/schedule # {enabled, cron, tz}
|
||||
GET /api/workflows/:id/schedule/preview?cron=…&tz=… # next 3 occurrences
|
||||
```
|
||||
|
||||
`PUT` validates the expression and the zone, then computes and stores
|
||||
`next_run_at`. The preview endpoint exists so the browser and the scheduler
|
||||
agree on what a cron string means — a client-side cron parser that disagrees
|
||||
with the server by one field is a bug found in production, at night.
|
||||
|
||||
### Frontend
|
||||
|
||||
- **Workflow page**: a schedule card with preset buttons (hourly, nightly at
|
||||
HH:MM, weekly on DAY at HH:MM) that write cron underneath, a raw cron field
|
||||
for anything else, a timezone select, and the next three occurrences rendered
|
||||
from the preview endpoint in mono.
|
||||
- **Workflows list**: a schedule chip and the next run as relative time.
|
||||
- **Skips are surfaced**, not just stored: a warning line reading
|
||||
"Skipped Sun 02:00 — previous run still active". Recording a reason nobody
|
||||
reads is the same as not recording one.
|
||||
|
||||
---
|
||||
|
||||
## Out of scope
|
||||
|
||||
**Notification on scheduled-run failure.** It needs the monitor channel
|
||||
machinery pointed at workflow outcomes and its own answer to what counts as
|
||||
failure — a non-zero exit on a step with `on_failure: continue` is not
|
||||
obviously an alert. Visibility in this change is the run list and the recorded
|
||||
skip reason. Excluded deliberately, not overlooked.
|
||||
|
||||
**Tag-scoped permissions.** Roles stay instance-wide. Tags describe servers;
|
||||
they do not yet gate who may act on them.
|
||||
|
||||
**Inventory-derived tags.** Reserved via the `sys:` prefix, not implemented.
|
||||
|
||||
---
|
||||
|
||||
## Migration and compatibility
|
||||
|
||||
No migration is required. `Tags`, `TargetTags` and `Schedule` are all
|
||||
`omitempty` and absent means what it meant before: no tags, no selector, no
|
||||
schedule. Existing workflows keep their explicit server lists and behave
|
||||
identically.
|
||||
|
||||
The wildcard tag index and the `next_run_at` index are declared by a new
|
||||
`EnsureServerIndexes`, following the convention `EnsureSecretIndexes` and
|
||||
`EnsureWorkflowIndexes` already set: it warns rather than aborting boot,
|
||||
because a missing index degrades
|
||||
tag filtering to a collection scan on a small collection rather than breaking
|
||||
the fleet list.
|
||||
|
||||
## Testing
|
||||
|
||||
- `ResolveTargets`: union deduplicates; AND across tag keys; empty/empty
|
||||
returns `ErrNoTargets`; offline servers are included.
|
||||
- Tag validation: charset, length caps, tag count cap, malformed `?tag=` is a
|
||||
400.
|
||||
- Schedule validation: bad cron and unknown zone both 400; `next_run_at` is
|
||||
computed in the stored zone, verified across a DST boundary.
|
||||
- Scheduler claim: two concurrent claims of the same due workflow start exactly
|
||||
one run.
|
||||
- Grace window: due 10 minutes ago runs; due 2 hours ago records `missed`.
|
||||
- Overlap: an active run yields `already_running` and no second run.
|
||||
- Preview endpoint and the scheduler agree on the next occurrence for a table
|
||||
of expressions, including a DST-crossing one.
|
||||
@@ -1,538 +0,0 @@
|
||||
# Package inventory and CVE findings
|
||||
|
||||
Date: 2026-08-06
|
||||
|
||||
Agents report the packages installed on each server. The control plane matches
|
||||
them against distro security feeds and raises findings that link straight to
|
||||
the patching path that already exists. A finding nobody can fix today can be
|
||||
accepted with a reason and an expiry date rather than sitting red forever.
|
||||
|
||||
This is one of four sub-projects sketched together and deliberately separated:
|
||||
|
||||
| # | Sub-project | Depends on |
|
||||
| - | ----------- | ---------- |
|
||||
| A | **Package inventory + CVE findings** — this spec | nothing |
|
||||
| B | Container/service registry | nothing |
|
||||
| C | Container image scanning | A and B |
|
||||
| D | Compliance profiles (baseline assertions) | shares A's findings UI only |
|
||||
|
||||
A and B are independent of one another. C is the joiner and must not be
|
||||
designed before both exist. D shares a page with A and nothing else — a
|
||||
different collector, a different evaluation model and a different remediation
|
||||
story — so folding it in here would double the size for no shared machinery.
|
||||
|
||||
Scope of this spec is **A, Linux only.** Windows needs a separate source
|
||||
(MSRC CVRF), a separate collector (`Get-HotFix` plus registry) and a KB
|
||||
supersedence matcher that shares no code with the Linux path. That matches the
|
||||
existing position that Windows agents are second-class by design, and the six
|
||||
package managers `updates.go` already detects cover the whole Linux surface.
|
||||
|
||||
---
|
||||
|
||||
## The trap this design is built around
|
||||
|
||||
Distributions **backport** security fixes without changing the upstream
|
||||
version. Ubuntu ships `openssl 3.0.2-0ubuntu1.15` patched against
|
||||
CVE-2023-0286; NVD says version 3.0.2 is vulnerable. Matching installed
|
||||
versions against NVD or CPE ranges therefore reports a fleet full of criticals
|
||||
that are all already fixed.
|
||||
|
||||
That is not merely noisy. It is fatal to the feature: once the first report is
|
||||
mostly wrong, nobody reads the second one, and a genuine finding is lost in the
|
||||
noise it created. Everything below follows from refusing to make that mistake.
|
||||
|
||||
The correct source is the **distribution's own security feed**, keyed on the
|
||||
distribution's own version string — Debian and Ubuntu OVAL/USN, Red Hat OVAL
|
||||
v2, Alpine secdb. `trivy-db` is those feeds pre-merged into one BoltDB
|
||||
artifact, rebuilt every six hours and published as an OCI artifact.
|
||||
|
||||
---
|
||||
|
||||
## Where the vulnerability data comes from
|
||||
|
||||
`trivy-db`, pulled server-side from `ghcr.io/aquasecurity/trivy-db:2`.
|
||||
|
||||
The alternative considered was querying OSV.dev per scan, which needs no
|
||||
storage and no puller. It was rejected on two counts: it requires outbound
|
||||
internet on every scan, which breaks air-gapped installs; and it sends the
|
||||
package list of a customer's entire fleet to a third party. The audience most
|
||||
likely to buy vulnerability scanning is the audience least willing to do that.
|
||||
|
||||
The blob is roughly 50MB, read-only, reproducible, and identified by a version
|
||||
number. **It is not stored in Mongo and not written to `/data`** —
|
||||
`server.persistence` defaults to off and nothing writes to `/data` any more.
|
||||
It does not need durable storage: whichever pod needs it pulls it to its own
|
||||
ephemeral temp directory. Nothing shared, nothing to back up, nothing to
|
||||
migrate.
|
||||
|
||||
`VANTAGE_TRIVY_DB_REF` overrides the default reference so a customer can mirror
|
||||
the artifact into their own registry. It also covers the anonymous ghcr rate
|
||||
limit, which the six-hourly pull cadence already makes unlikely to bite.
|
||||
|
||||
---
|
||||
|
||||
## Only the leader matches
|
||||
|
||||
This is the crux, and it falls out of the replica model already in the
|
||||
codebase.
|
||||
|
||||
Two things trigger matching, and they happen on different pods:
|
||||
|
||||
1. a fleet-wide rescan when `trivy-db` updates — naturally the leader's job
|
||||
2. a server's package list changing — handled by whichever pod holds *that
|
||||
agent's* command stream
|
||||
|
||||
If (2) matched inline, **every replica would need the 50MB database resident**,
|
||||
and a database refresh would have N pods racing to rescan the same fleet and N
|
||||
digests reaching the customer. That is the exact failure `RunAsLeader` exists
|
||||
to prevent, and it is the same argument that put `monitorsched` behind the
|
||||
lock.
|
||||
|
||||
So `ReportPackages` does not match. It upserts the package list and sets
|
||||
`scan_pending: true`. That is all it does.
|
||||
|
||||
`server/internal/vulnsched` then runs inside the **existing**
|
||||
`bus.RunAsLeader("housekeeping", …)` alongside `monitorsched`,
|
||||
`workflowsched` and the sweepers — one role, one lock. Every 60 seconds it:
|
||||
|
||||
1. pulls `trivy-db` if the local copy is older than six hours
|
||||
2. if the pulled version differs from `vulndb_meta.db_version`, marks **every**
|
||||
server `scan_pending`
|
||||
3. matches all `scan_pending` servers, clears the flag, diffs against existing
|
||||
findings
|
||||
4. emits **one** digest per tick covering everything newly opened
|
||||
|
||||
Step 4 is why batching is structural rather than bolted on. A `trivy-db`
|
||||
refresh can open several hundred findings across a fleet at once; one message
|
||||
per finding would rate-limit the webhook or get the channel muted, and either
|
||||
way the customer stops receiving the alerts they are paying for. The tick is
|
||||
already the natural batch boundary, so **the failure cannot occur by
|
||||
construction** rather than by a debounce someone has to maintain.
|
||||
|
||||
`scan_pending` lives on the document rather than in memory, for the same reason
|
||||
`next_run_at` and `workflow_log_seq` do: a leader handover between marking and
|
||||
scanning would otherwise lose it. A handover costs the new leader one re-pull
|
||||
of the database.
|
||||
|
||||
The cost of this indirection is up to 60 seconds between an agent reporting a
|
||||
changed package set and its findings updating. For vulnerability data that is
|
||||
nothing, and it buys a single matching path instead of two.
|
||||
|
||||
---
|
||||
|
||||
## Components
|
||||
|
||||
```
|
||||
agent/internal/packages/ collect installed packages + /etc/os-release
|
||||
proto/ ReportPackages RPC
|
||||
server/internal/vulndb/ puller, BoltDB access, matcher
|
||||
server/internal/vulnsched/ leader-owned tick: pull, scan, digest
|
||||
server/internal/services/ findings, acceptance, alert rules
|
||||
web/app/(app)/vulnerabilities/ fleet board; plus two server-detail tabs
|
||||
```
|
||||
|
||||
`vulnsched` takes the dependencies it needs — `LogEvent` and the notification
|
||||
dispatch — as a `vulnsched.Deps` injected from `main.go`, following
|
||||
`workflowsched`. The manual rescan endpoint does not call into `vulnsched` at
|
||||
all: it sets `scan_pending` on every server and lets the next tick find them,
|
||||
so there is no path by which `services` imports the scheduler and no cycle to
|
||||
avoid later.
|
||||
|
||||
---
|
||||
|
||||
## The wire path
|
||||
|
||||
A new `ReportPackages` RPC on the agent's existing hourly loop — the same
|
||||
`runUpdateCheck` cadence, reusing `updates.go`'s `detectPM()`.
|
||||
|
||||
```protobuf
|
||||
rpc ReportPackages(ReportPackagesRequest) returns (ReportPackagesResponse);
|
||||
|
||||
message ReportPackagesRequest {
|
||||
string server_id = 1;
|
||||
string agent_token = 2;
|
||||
string hash = 3; // sha256 of the sorted list
|
||||
OSRelease os = 4;
|
||||
repeated InstalledPackage packages = 5; // omitted when only offering a hash
|
||||
}
|
||||
|
||||
message ReportPackagesResponse {
|
||||
bool need_full = 1; // hash differs; resend with packages populated
|
||||
}
|
||||
```
|
||||
|
||||
The agent calls once with `packages` empty. `need_full` true means the hash
|
||||
differs from what the server holds, and the agent immediately calls again with
|
||||
the list populated.
|
||||
|
||||
The agent sends a SHA-256 of its sorted package list first. If it matches what
|
||||
the server already holds, the server answers `unchanged` and the ~150KB body is
|
||||
never sent. A machine's package set changes rarely, so almost every hour costs
|
||||
one small message, and the rare changed hour costs one extra round trip.
|
||||
|
||||
Folding the list into the existing 15-minute `InventoryReport` static snapshot
|
||||
was rejected: it would re-send ~150KB per server every 15 minutes regardless of
|
||||
change, roughly 40MB/hour of gRPC traffic on a 100-server fleet to transmit
|
||||
data that is almost always identical.
|
||||
|
||||
---
|
||||
|
||||
## Data model
|
||||
|
||||
Four new collections. Every one carries `instance_id` except `vulndb_meta`,
|
||||
which is explained below.
|
||||
|
||||
### `server_packages` — one document per server, not per package
|
||||
|
||||
```go
|
||||
type ServerPackages struct {
|
||||
ID primitive.ObjectID `bson:"_id"`
|
||||
InstanceID primitive.ObjectID `bson:"instance_id"`
|
||||
ServerID string `bson:"server_id"`
|
||||
OS OSRelease `bson:"os"` // family, version_id, arch
|
||||
Hash string `bson:"hash"` // sha256 of the sorted list
|
||||
Packages []InstalledPackage `bson:"packages"`
|
||||
CollectedAt time.Time `bson:"collected_at"`
|
||||
ScanPending bool `bson:"scan_pending"`
|
||||
ScannedAt time.Time `bson:"scanned_at"`
|
||||
Status string `bson:"status"` // ok | unsupported
|
||||
DBVersion int `bson:"db_version"` // last matched against
|
||||
}
|
||||
|
||||
type InstalledPackage struct {
|
||||
Name string `bson:"name"`
|
||||
Version string `bson:"version"` // distro version string, verbatim
|
||||
Epoch int `bson:"epoch,omitempty"`
|
||||
Arch string `bson:"arch"`
|
||||
SourceName string `bson:"source_name,omitempty"`
|
||||
}
|
||||
```
|
||||
|
||||
One document rather than two thousand is what makes a report a **single atomic
|
||||
upsert with no delta logic** — the hash already established that something
|
||||
changed, so there is nothing to reconcile field by field. A typical Linux host
|
||||
lands near 150KB, comfortably inside the 16MB document limit.
|
||||
|
||||
Indexes: `{instance_id, server_id}` unique, and a multikey
|
||||
`{instance_id, "packages.name"}` for fleet-wide package search.
|
||||
|
||||
`SourceName` is not decoration. **Debian and Ubuntu advisories are keyed on the
|
||||
source package**: a CVE against `openssl` covers the binaries `libssl3`,
|
||||
`openssl` and `libssl-dev`, so matching on binary name alone misses two of the
|
||||
three.
|
||||
|
||||
`OS.VersionID` selects the feed. Ubuntu 22.04 and 24.04 publish different fixed
|
||||
versions for the same CVE, so a scan without it is guesswork.
|
||||
|
||||
### `vuln_findings` — one document per (server, CVE, package)
|
||||
|
||||
```go
|
||||
type VulnFinding struct {
|
||||
ID primitive.ObjectID `bson:"_id"`
|
||||
InstanceID primitive.ObjectID `bson:"instance_id"`
|
||||
ServerID string `bson:"server_id"`
|
||||
|
||||
CVEID string `bson:"cve_id"`
|
||||
PackageName string `bson:"package_name"`
|
||||
Installed string `bson:"installed_version"`
|
||||
FixedIn string `bson:"fixed_in,omitempty"`
|
||||
Severity string `bson:"severity"`
|
||||
CVSSScore float64 `bson:"cvss_score,omitempty"`
|
||||
Title string `bson:"title,omitempty"`
|
||||
References []string `bson:"references,omitempty"`
|
||||
|
||||
State string `bson:"state"` // open | fixed | accepted
|
||||
FirstSeen time.Time `bson:"first_seen"`
|
||||
LastSeen time.Time `bson:"last_seen"`
|
||||
FixedAt *time.Time `bson:"fixed_at,omitempty"`
|
||||
Accepted *Acceptance `bson:"accepted,omitempty"`
|
||||
}
|
||||
|
||||
type Acceptance struct {
|
||||
By primitive.ObjectID `bson:"by"`
|
||||
Reason string `bson:"reason"`
|
||||
Until time.Time `bson:"until"`
|
||||
At time.Time `bson:"at"`
|
||||
}
|
||||
```
|
||||
|
||||
Unique on `{instance_id, server_id, cve_id, package_name}`. That key is what
|
||||
makes a rescan an idempotent upsert rather than a duplicate factory, and it is
|
||||
what lets `first_seen` survive across scans. Query index
|
||||
`{instance_id, state, severity}`.
|
||||
|
||||
**An empty `FixedIn` is a real and common state** and must never be conflated
|
||||
with "not vulnerable". A CVE with no vendor fix published yet is exactly the
|
||||
finding people most need to see, and also the one that most needs acceptance,
|
||||
because there is nothing to patch.
|
||||
|
||||
Findings are **not deleted when a package is patched**. State moves to `fixed`
|
||||
with `fixed_at` set, so "what did we remediate last quarter" remains
|
||||
answerable — which is the question an auditor asks.
|
||||
|
||||
### `vulndb_meta` — singleton, deliberately unscoped
|
||||
|
||||
`db_version`, `pulled_at`, `last_full_scan_at`, `last_error`. It carries no
|
||||
`instance_id` because the vulnerability database is a property of the
|
||||
deployment, not of a tenant. Same reasoning as `migrations`.
|
||||
|
||||
### `vuln_alert_rules`
|
||||
|
||||
`instance_id`, `name`, `enabled`, `min_severity`, `tags map[string]string`,
|
||||
`channel_ids []`, timestamps.
|
||||
|
||||
The tag filter resolves through **`services.ResolveTargets`**, not a second
|
||||
matcher. That function is already the single answer to which servers a
|
||||
selector touches, and an alert rule that disagreed with a workflow about what
|
||||
`env:prod` means would be worse than having no filter at all.
|
||||
|
||||
---
|
||||
|
||||
## The matching engine
|
||||
|
||||
```
|
||||
server/internal/vulndb/
|
||||
pull.go OCI fetch → temp dir, version compare against vulndb_meta
|
||||
db.go BoltDB open, advisory lookup by (ecosystem, source, version)
|
||||
match.go per-family matching, severity resolution
|
||||
version.go dispatch to deb/rpm/apk comparator by OS family
|
||||
```
|
||||
|
||||
Dependencies: `github.com/aquasecurity/trivy-db` for the BoltDB schema, plus
|
||||
`go-deb-version`, `go-rpm-version` and `go-apk-version` — each a small
|
||||
standalone module doing one job. The roughly 200 lines of per-distro advisory
|
||||
lookup are ours.
|
||||
|
||||
Importing `trivy` itself was rejected: it would pull a very large transitive
|
||||
dependency tree into the server binary for one feature, and its Go API carries
|
||||
no stability guarantee across minor versions. Shelling out to the `trivy`
|
||||
binary against a generated SBOM was rejected for shipping a second binary in
|
||||
the image and turning a library call into subprocess lifecycle, timeouts and
|
||||
output-format drift.
|
||||
|
||||
### Why the comparators are bought rather than written
|
||||
|
||||
Version ordering is where this feature lives or dies, and its failure mode is
|
||||
silent. `dpkg` ordering has epochs, and `~` sorts *before* the empty string, so
|
||||
`3.0.2-0ubuntu1.15~rc1` precedes `3.0.2-0ubuntu1.15`. `rpmvercmp` has its own
|
||||
segment rules and treats `~` and `^` differently again. A `strings.Compare` or
|
||||
a semver parse orders `1.9` above `1.10` and reports a vulnerable fleet as
|
||||
clean — a false negative, which nobody notices until it matters.
|
||||
|
||||
### Scanning one server
|
||||
|
||||
1. Load `server_packages`; resolve OS family and version to a `trivy-db`
|
||||
ecosystem.
|
||||
2. **Unsupported ecosystem → record `status: unsupported`, clear the flag,
|
||||
write no findings.**
|
||||
3. For each package: resolve source name, look up advisories, compare versions.
|
||||
4. Upsert vulnerable results as `open`, preserving `first_seen`.
|
||||
5. Any currently-`open` finding absent from this result set → `fixed`, stamp
|
||||
`fixed_at`.
|
||||
6. Any `accepted` finding past its `until` → back to `open`.
|
||||
7. Clear `scan_pending`, stamp `scanned_at` and `db_version`.
|
||||
|
||||
Steps 5 and 6 must run in that order, so a finding that is both absent and
|
||||
expired settles as `fixed` rather than reopening on a package that no longer
|
||||
carries it.
|
||||
|
||||
Step 2 matters as much as any of the matching. Arch has no feed in `trivy-db`,
|
||||
so an Arch host must report **unsupported**, never "0 findings". Reporting
|
||||
clean when the truth is unknown is the same class of lie as a silently stale
|
||||
database, and it is the reason `vulndb_meta.pulled_at` appears on screen rather
|
||||
than only in a log.
|
||||
|
||||
### Severity
|
||||
|
||||
Resolved **vendor → NVD → unknown**, in that order, never invented.
|
||||
|
||||
This will surface as "why is this critical CVE marked low", and the answer is
|
||||
that Debian and Red Hat routinely downgrade an NVD score because the vulnerable
|
||||
code path is not reachable in their build. Their rating is the accurate one for
|
||||
that package, and showing NVD's above it would manufacture work that does not
|
||||
need doing.
|
||||
|
||||
---
|
||||
|
||||
## Findings lifecycle
|
||||
|
||||
`open | fixed | accepted`.
|
||||
|
||||
An accepted finding is suppressed from counts and alerts until its `until`
|
||||
date, then reopens automatically. A reason is required.
|
||||
|
||||
Acceptance with a mandatory expiry, rather than permanent dismissal, is what
|
||||
keeps the feature usable in both directions. Without any acceptance mechanism,
|
||||
a kernel CVE awaiting a reboot window sits red indefinitely and trains people
|
||||
to ignore the page. With permanent dismissal, accepted findings accumulate
|
||||
silently and nobody revisits them — the dismissal list becomes where risk goes
|
||||
to be forgotten, which is precisely what an auditor asks to see.
|
||||
|
||||
Retention: `settings.vuln_finding_retention_days`, a `*int` on the same pattern
|
||||
as `workflow_log_retention_days` — nil means 90 days, 0 means forever. Only
|
||||
`fixed` findings are swept, by a `StartVulnSweeper` inside the same
|
||||
`RunAsLeader("housekeeping", …)` as the existing sweepers. `open` and
|
||||
`accepted` findings are never swept at any setting.
|
||||
|
||||
---
|
||||
|
||||
## Alerting
|
||||
|
||||
Per-org rules over the existing `notification_channels`: severity threshold,
|
||||
optional tag filter, target channels.
|
||||
|
||||
A rescan emits one message summarising what newly opened — "12 new critical
|
||||
across 4 servers" — never one message per finding. See the leader section for
|
||||
why the tick boundary makes this structural.
|
||||
|
||||
Modelling findings as a monitor type was rejected. It would reuse monitors'
|
||||
state machine and channel wiring for free, but monitors are up/down for one
|
||||
endpoint with retries and hourly rollups, none of which means anything for a
|
||||
CVE; most fields would be disabled in the UI and the uptime graphs would be
|
||||
polluted with a signal that is not uptime.
|
||||
|
||||
This adds one `notify` payload type and a `vuln_digest.html.tmpl` /
|
||||
`vuln_digest.txt.tmpl` pair in `shared/mail`. Note that `shared/mail` templates
|
||||
are parsed in `init()`, so a mistyped field is a boot-time panic — CLAUDE.md
|
||||
describes a `render_test.go` guarding against exactly this, but **that file does
|
||||
not exist**; the repository has no Go tests at all, and by instruction this
|
||||
feature adds none. The template pair must therefore be verified by starting the
|
||||
binary and sending one digest through a real channel.
|
||||
|
||||
---
|
||||
|
||||
## Entitlement
|
||||
|
||||
The feature name is `vuln_scanning`, and it crosses the two services the way
|
||||
every other feature does:
|
||||
|
||||
- **admin** carries it as a per-instance entitlement toggle, so it can later be
|
||||
priced as a catalogue `feature` component without a second migration;
|
||||
- **the licence** snapshots it into `License.Features []string` at issue time;
|
||||
- **the server** asks `lic.HasFeature("vuln_scanning")` and never switches on
|
||||
tier, so changing what a tier includes needs no server release.
|
||||
|
||||
Off on Free.
|
||||
|
||||
**The gate is checked at `ReportPackages`, not at display.** Gating only the UI
|
||||
would still pay every write cost, and storage is the expensive half.
|
||||
|
||||
The agent learns of it through the existing 30-second `SyncKeys` poll:
|
||||
`SyncResponse` gains a `collect_packages` bool, and the hourly loop skips
|
||||
collection entirely when it is false. So an ungated instance produces no
|
||||
collection, no gRPC body, no document and no storage. `ReportPackages` still
|
||||
re-checks the entitlement server-side and refuses — the agent flag is an
|
||||
optimisation, the server check is the boundary.
|
||||
|
||||
Turning the feature off does not delete existing findings; they stop being
|
||||
served and stop updating. Deletion is the instance-deletion path's job.
|
||||
|
||||
---
|
||||
|
||||
## REST API
|
||||
|
||||
```
|
||||
GET /api/vulnerabilities # filter: severity, state, server, tags
|
||||
GET /api/vulnerabilities/summary # severity counts + database freshness
|
||||
POST /api/vulnerabilities/rescan # marks all scan_pending (owner|admin)
|
||||
POST /api/vulnerabilities/:id/accept # reason + until (owner|admin)
|
||||
DELETE /api/vulnerabilities/:id/accept # (owner|admin)
|
||||
GET /api/servers/:id/vulnerabilities
|
||||
GET /api/servers/:id/packages
|
||||
GET /api/packages/search?name= # fleet-wide
|
||||
GET,POST /api/vuln-rules · PUT,DELETE /api/vuln-rules/:id
|
||||
```
|
||||
|
||||
Every mutating path writes an audit event, as all of them do. Acceptance is the
|
||||
one decision people will be asked to justify, so `by`, `reason`, `until` and
|
||||
`at` land in the audit record and not only on the document.
|
||||
|
||||
---
|
||||
|
||||
## UI
|
||||
|
||||
`/vulnerabilities` is a fleet board **grouped by CVE** — one row per CVE with
|
||||
an affected-server count, expandable to the individual servers. The same CVE
|
||||
across 40 servers is one decision, and a flat list of findings makes it look
|
||||
like forty.
|
||||
|
||||
Server detail gains **Vulnerabilities** and **Packages** tabs. Alert rules go
|
||||
on `/settings/notifications`, beside the channels they consume.
|
||||
|
||||
Remediation introduces no new mechanism: a finding carrying `fixed_in` renders
|
||||
an **Apply updates** action calling the existing
|
||||
`POST /api/servers/:id/apply-updates`, which is already `ApplyUpdatesCmd`. See
|
||||
it, patch it, one place — and no second patching path to keep consistent with
|
||||
the first.
|
||||
|
||||
Database freshness is shown wherever findings are, not tucked into settings. A
|
||||
fleet scanning against a three-week-old database must say so rather than
|
||||
quietly report all-clear.
|
||||
|
||||
---
|
||||
|
||||
## Verification
|
||||
|
||||
**No automated tests.** The repository has none today, and by explicit
|
||||
instruction this feature adds none — no `*_test.go`, no frontend test files.
|
||||
That is a deliberate decision by the repository owner, recorded here so the
|
||||
absence reads as a choice rather than an omission.
|
||||
|
||||
It does change the risk profile, and the places it changes it are worth naming,
|
||||
because each fails by producing a **wrong answer rather than a crash**:
|
||||
|
||||
- **Version comparison.** The backport case — installed `1:3.0.2-0ubuntu1.15`
|
||||
against advisory fixed-in `1:3.0.2-0ubuntu1.15` resolving to *not
|
||||
vulnerable* — plus tilde ordering (`1.0~rc1` < `1.0`), epoch dominance
|
||||
(`1:1.0` > `2.0`) and `1.9` < `1.10`. Wrong here means a vulnerable fleet
|
||||
reported clean.
|
||||
- **Source-package fan-out.** One advisory against `openssl` must flag
|
||||
`libssl3`, `openssl` and `libssl-dev`. Matching on binary name alone silently
|
||||
finds one of three.
|
||||
- **`first_seen` preservation.** An upsert that overwrites it makes every
|
||||
finding look discovered today, and nothing surfaces that until someone reads
|
||||
a report.
|
||||
- **Fixed-before-reopen ordering.** A finding both absent from a scan and past
|
||||
its acceptance expiry must settle `fixed`, not reopen.
|
||||
|
||||
The implementation plan carries a manual verification table for each, to be
|
||||
walked before the relevant task is committed. They are the substitute for the
|
||||
tests, not a formality.
|
||||
|
||||
---
|
||||
|
||||
## Failure modes
|
||||
|
||||
| Failure | Behaviour |
|
||||
| ------- | --------- |
|
||||
| Database pull fails | Keep the last good copy and serve stale. Record `last_error`, surface `pulled_at` age. **Never clear findings** — a network blip must not read as "all fixed" |
|
||||
| Unsupported distribution | `status: unsupported`, not zero findings |
|
||||
| Agent stops reporting | Findings persist and `collected_at` age is shown. No auto-expiry: a silent agent is not a patched server |
|
||||
| `trivy-db` schema version bumps | The puller refuses an unknown schema rather than mis-parsing it |
|
||||
| ghcr anonymous rate limit | Backoff; `VANTAGE_TRIVY_DB_REF` mirrors to a private registry |
|
||||
| Leadership lost mid-scan | The context is cancelled and the scan returns; `scan_pending` is still set, so the next leader picks it up |
|
||||
| Instance deleted | **`server_packages` and `vuln_findings` must be added to the control plane's instance-deletion collection list.** Easy to miss, and missing it orphans a tenant's package data indefinitely |
|
||||
|
||||
---
|
||||
|
||||
## Environment variables
|
||||
|
||||
| Name | Required | Notes |
|
||||
| ---- | -------- | ----- |
|
||||
| `VANTAGE_TRIVY_DB_REF` | no | default `ghcr.io/aquasecurity/trivy-db:2`. Point at a mirror for air-gapped installs or to avoid the anonymous ghcr rate limit |
|
||||
| `VANTAGE_VULNDB_DISABLED` | no | disables the puller and the scheduler entirely. Findings already written are still served and still marked stale |
|
||||
|
||||
---
|
||||
|
||||
## Deliberately out of scope
|
||||
|
||||
- **Windows.** Separate source, collector and matcher; its own spec.
|
||||
- **Container image scanning.** Sub-project C; needs the container registry.
|
||||
- **Compliance baseline assertions.** Sub-project D; shares this findings UI
|
||||
and nothing else.
|
||||
- **Language-level dependency scanning** (npm, pip, Go modules). `trivy-db`
|
||||
covers these ecosystems, but finding the manifests on a host is a different
|
||||
collection problem from asking the package manager what is installed.
|
||||
- **Automatic patching on a finding.** Remediation is one click, not zero. An
|
||||
unattended upgrade triggered by a CVE feed is a fleet-wide change driven by a
|
||||
third party's data, which is not a decision to take away from an operator.
|
||||
@@ -1,441 +0,0 @@
|
||||
# Workload registry
|
||||
|
||||
Date: 2026-08-06
|
||||
|
||||
Agents enumerate what each server actually runs — Docker containers, the
|
||||
compose stacks grouping them, and systemd services — and report it to the
|
||||
control plane. Containers and units can be started, stopped and restarted from
|
||||
the UI, and a bounded snapshot of their logs can be read without opening a
|
||||
console.
|
||||
|
||||
This is **sub-project B** of the four sketched in
|
||||
`2026-08-06-package-inventory-and-cve-findings-design.md`:
|
||||
|
||||
| # | Sub-project | Depends on |
|
||||
| - | ----------- | ---------- |
|
||||
| A | Package inventory + CVE findings — its own spec | nothing |
|
||||
| B | **Workload registry** — this spec | nothing |
|
||||
| C | Container image scanning | A and B |
|
||||
| D | Compliance profiles | shares A's findings UI only |
|
||||
|
||||
A and B are independent. C is the joiner and must not be designed before both
|
||||
exist: it needs B's image list and A's findings model.
|
||||
|
||||
**Workload** is the domain word throughout: one container or one systemd unit.
|
||||
It gives the collection, the commands and the page a single honest name rather
|
||||
than saying "container or service" in every identifier.
|
||||
|
||||
Scope is **Linux only**, matching sub-project A and the existing position that
|
||||
Windows agents are second-class by design. Docker runs on Windows; systemd does
|
||||
not, and half a feature per platform is worse than a clear line.
|
||||
|
||||
---
|
||||
|
||||
## What this is for
|
||||
|
||||
The control plane can manage a fleet's keys, run workflows across it and watch
|
||||
its endpoints, but it has no idea what any of those servers actually *runs*.
|
||||
"Restart nginx on that box" means opening a console. "Which of these 80 servers
|
||||
is still on the old image" is unanswerable.
|
||||
|
||||
---
|
||||
|
||||
## Reporting and refresh are one path
|
||||
|
||||
The agent reports on its own 60-second ticker through a `ReportWorkloads` RPC,
|
||||
using the same hash short-circuit as the package report: it offers a SHA-256 of
|
||||
the sorted workload list, and sends the body only when the server does not
|
||||
already hold that hash. An unchanged list costs one small message, which on a
|
||||
60-second cadence is the common case by a wide margin.
|
||||
|
||||
The on-demand refresh **does not return data**. `RefreshWorkloadsCmd` carries
|
||||
no payload back; it makes the agent report immediately through the normal RPC,
|
||||
and the UI refetches the stored document.
|
||||
|
||||
That is deliberate. A refresh that returned workloads inline would be a second
|
||||
writer for the same collection, arriving by a different route, with its own
|
||||
serialisation and its own opportunity to disagree with the periodic one. One
|
||||
writer, one shape; the refresh is a nudge, not a channel.
|
||||
|
||||
Opening a server's Workloads tab dispatches a refresh, so what is on screen is
|
||||
live rather than up to a minute stale. That matters because the page has a
|
||||
Restart button on it: a stale list is not merely a wrong impression, it is a
|
||||
wrong action aimed at a container that already died.
|
||||
|
||||
## What does answer back
|
||||
|
||||
Two operations genuinely return something:
|
||||
|
||||
| Command | Answers with |
|
||||
| ------- | ------------ |
|
||||
| `ControlWorkloadCmd{kind, id, action}` | the existing `CommandResult` — ok or error |
|
||||
| `WorkloadLogsCmd{kind, id, tail}` | a new `WorkloadLogsResult{command_id, text, truncated}` |
|
||||
|
||||
Both ride the proven path: `commandDispatcher.send()` for request and ack, and
|
||||
a `WorkloadResults` registry mirroring `StepResults.Await`/`Deliver` over the
|
||||
bus. **`Await` must subscribe before the command is dispatched** — the pod
|
||||
driving the request is usually not the pod holding the agent's stream, and a
|
||||
fast agent otherwise answers into a channel nobody has joined. This is not a
|
||||
new hazard; it is the one `stepresults.go` already documents.
|
||||
|
||||
```protobuf
|
||||
rpc ReportWorkloads(ReportWorkloadsRequest) returns (ReportWorkloadsResponse);
|
||||
|
||||
message ReportWorkloadsRequest {
|
||||
string server_id = 1;
|
||||
string agent_token = 2;
|
||||
string hash = 3;
|
||||
bool docker_ok = 4;
|
||||
string docker_error = 5;
|
||||
bool systemd_ok = 6;
|
||||
string systemd_error = 7;
|
||||
repeated Workload workloads = 8; // empty on the offer call
|
||||
}
|
||||
|
||||
message ReportWorkloadsResponse {
|
||||
bool need_full = 1;
|
||||
}
|
||||
|
||||
// ServerCommand gains three variants.
|
||||
message RefreshWorkloadsCmd {}
|
||||
|
||||
message ControlWorkloadCmd {
|
||||
string kind = 1; // "container" | "unit"
|
||||
string id = 2;
|
||||
string action = 3; // "start" | "stop" | "restart"
|
||||
}
|
||||
|
||||
message WorkloadLogsCmd {
|
||||
string kind = 1;
|
||||
string id = 2;
|
||||
int32 tail = 3;
|
||||
}
|
||||
|
||||
// AgentMessage gains one variant.
|
||||
message WorkloadLogsResult {
|
||||
string command_id = 1;
|
||||
string text = 2;
|
||||
bool truncated = 3;
|
||||
string error = 4;
|
||||
}
|
||||
```
|
||||
|
||||
The offer-then-send handshake is the package report's, unchanged: the agent
|
||||
calls once with `workloads` empty, and resends with the body only if the
|
||||
response sets `need_full`.
|
||||
|
||||
An agent whose stream no pod holds gets a 503 from the dispatcher, as
|
||||
everything else does. Commands are not queued: a command whose owner died must
|
||||
fail loudly rather than be delivered to nobody while the operator is told it
|
||||
worked.
|
||||
|
||||
---
|
||||
|
||||
## Not gated by licence
|
||||
|
||||
Unlike CVE scanning, this reads as core fleet management rather than a premium
|
||||
add-on, so v1 ships to every instance with no entitlement check.
|
||||
|
||||
If that changes it is a one-line `HasFeature` check at `ReportWorkloads`,
|
||||
gating collection rather than display — the same placement and the same
|
||||
reasoning as sub-project A, where gating the UI alone would still pay every
|
||||
write cost.
|
||||
|
||||
---
|
||||
|
||||
## Data model
|
||||
|
||||
One new collection, `server_workloads`, one document per server, mirroring
|
||||
`server_packages`.
|
||||
|
||||
```go
|
||||
type ServerWorkloads struct {
|
||||
ID bson.ObjectID `bson:"_id,omitempty" json:"-"`
|
||||
InstanceID string `bson:"instance_id" json:"-"`
|
||||
ServerID string `bson:"server_id" json:"server_id"`
|
||||
Hash string `bson:"hash" json:"hash"`
|
||||
Workloads []Workload `bson:"workloads" json:"workloads"`
|
||||
CollectedAt time.Time `bson:"collected_at" json:"collected_at"`
|
||||
|
||||
DockerOK bool `bson:"docker_ok" json:"docker_ok"`
|
||||
DockerError string `bson:"docker_error,omitempty" json:"docker_error,omitempty"`
|
||||
SystemdOK bool `bson:"systemd_ok" json:"systemd_ok"`
|
||||
SystemdError string `bson:"systemd_error,omitempty" json:"systemd_error,omitempty"`
|
||||
}
|
||||
|
||||
type Workload struct {
|
||||
Kind string `bson:"kind" json:"kind"` // "container" | "unit"
|
||||
ID string `bson:"id" json:"id"` // container id, or unit name
|
||||
Name string `bson:"name" json:"name"`
|
||||
State string `bson:"state" json:"state"`
|
||||
Health string `bson:"health,omitempty" json:"health,omitempty"`
|
||||
Image string `bson:"image,omitempty" json:"image,omitempty"`
|
||||
Stack string `bson:"stack,omitempty" json:"stack,omitempty"`
|
||||
Ports []string `bson:"ports,omitempty" json:"ports,omitempty"`
|
||||
Restarts int `bson:"restarts,omitempty" json:"restarts,omitempty"`
|
||||
StartedAt time.Time `bson:"started_at,omitempty" json:"started_at,omitempty"`
|
||||
Protected bool `bson:"protected" json:"protected"`
|
||||
}
|
||||
```
|
||||
|
||||
`State` is normalised across the two kinds: containers report `running`,
|
||||
`exited`, `paused`, `restarting`, `created`; units report `active`, `inactive`,
|
||||
`failed`, `activating`. They are deliberately **not** collapsed into a shared
|
||||
vocabulary — a failed unit and an exited container mean different things, and
|
||||
flattening them would lose the distinction the operator needs.
|
||||
|
||||
Indexes: `{instance_id, server_id}` unique, plus multikey
|
||||
`{instance_id, "workloads.image"}` for the fleet-wide "which servers run image
|
||||
X" query.
|
||||
|
||||
### Why the OK/Error pairs exist
|
||||
|
||||
A host with no Docker installed and a host where Docker is installed and
|
||||
running nothing both produce an empty list. One should read "not in use here",
|
||||
the other "nothing running", and only the second deserves any alarm.
|
||||
|
||||
The error strings separate a third case the booleans alone cannot: Docker
|
||||
installed with the daemon down. "Not installed" and "installed but not
|
||||
responding" are different problems with different fixes, and collapsing them
|
||||
into one false boolean throws away the only thing that tells them apart.
|
||||
|
||||
### Why `Protected` is reported rather than derived
|
||||
|
||||
The agent already knows which unit and container it is. Sending that up lets
|
||||
the UI render the action disabled with a reason instead of offering a button
|
||||
whose refusal is already known.
|
||||
|
||||
The field is the courtesy; the agent's own check is the boundary. See the
|
||||
control section.
|
||||
|
||||
### No history
|
||||
|
||||
A workload list is state, not a record. Nobody asks what containers ran last
|
||||
Tuesday, and keeping it would grow a collection per server per minute in
|
||||
exchange for a question nobody has.
|
||||
|
||||
---
|
||||
|
||||
## Collectors
|
||||
|
||||
### Docker: two commands, no English parsing
|
||||
|
||||
```
|
||||
docker ps -aq
|
||||
docker inspect --format '{{json .}}' <ids…>
|
||||
```
|
||||
|
||||
Not `docker ps --format '{{json .}}'` alone. That reports health and uptime
|
||||
inside a human `Status` string — `"Up 2 hours (healthy)"` — and anything built
|
||||
on it is parsing English that is localised, reworded between releases, and
|
||||
silently different for a paused or restarting container. `inspect` returns
|
||||
`State.Health.Status`, `State.StartedAt` and `RestartCount` as typed fields.
|
||||
Two execs instead of one, and no parser to be wrong.
|
||||
|
||||
`RestartCount` justifies the second call by itself: a container cycling is the
|
||||
single thing this page most needs to show, and it is invisible in a list that
|
||||
only ever says "Up".
|
||||
|
||||
Compose stacks come from the `com.docker.compose.project` label. **No YAML is
|
||||
read from disk** — the label is what Docker itself treats as authoritative, and
|
||||
a compose file on disk may not be what is actually running.
|
||||
|
||||
Docker absent, or a socket that cannot be reached, sets `DockerOK: false`. It
|
||||
is not an error and produces no log line: most servers in a fleet built around
|
||||
SSH key management will not have Docker, and treating the normal case as a
|
||||
fault makes the feature look broken on the majority of the estate.
|
||||
|
||||
### systemd: filtered on purpose
|
||||
|
||||
```
|
||||
systemctl list-units --type=service --state=running,failed --no-legend --plain --no-pager
|
||||
systemctl list-unit-files --type=service --state=enabled --no-legend --plain --no-pager
|
||||
```
|
||||
|
||||
Two calls because "running or failed" and "enabled but stopped" are different
|
||||
questions, and an enabled unit that is not running is exactly the one worth
|
||||
seeing.
|
||||
|
||||
Excluded by prefix: `systemd-`, `user@`, `session-`, `init.scope`. A typical
|
||||
host carries 300+ units, the platform's own accounting for most of them.
|
||||
Listing all of them buries the ten anyone cares about — the same failure mode
|
||||
as an unfiltered vulnerability report, and the same fix.
|
||||
|
||||
Column output rather than `--output=json`: the JSON flag requires systemd 246+,
|
||||
and this fleet includes older stable distributions. The column format has been
|
||||
stable considerably longer than the JSON one has existed.
|
||||
|
||||
---
|
||||
|
||||
## Control actions
|
||||
|
||||
```
|
||||
container: docker {start|stop|restart} <id>
|
||||
unit: systemctl {start|stop|restart} <unit>
|
||||
```
|
||||
|
||||
Owner or admin only. Every action writes an audit event naming the actor, the
|
||||
server and the target.
|
||||
|
||||
### The protected set
|
||||
|
||||
Computed agent-side: `vantage-agent.service`, plus the container ID read from
|
||||
`/proc/self/cgroup` should the agent ever be run inside a container.
|
||||
|
||||
The agent refuses those before doing anything. As with the console relay
|
||||
hardcoding `127.0.0.1` agent-side, **the control plane may name a target, but
|
||||
the agent decides what it will do to itself**. A server-side denylist alone
|
||||
would be bypassed by the next dispatch path someone adds, and the failure is
|
||||
unrecoverable from the UI: a server that stops its own agent goes offline, and
|
||||
the way back is SSH or physical access — precisely what this feature exists to
|
||||
avoid needing.
|
||||
|
||||
### Timeouts
|
||||
|
||||
`docker stop` waits on a container that may ignore SIGTERM. `systemctl stop`
|
||||
on a unit with a long `TimeoutStopSec` blocks for exactly as long as that says.
|
||||
Both run under a 90-second context, and a timeout returns a real error rather
|
||||
than an ack implying success.
|
||||
|
||||
---
|
||||
|
||||
## Logs
|
||||
|
||||
```
|
||||
container: docker logs --tail 500 --timestamps <id>
|
||||
unit: journalctl -u <unit> -n 500 --no-pager --output=short-iso
|
||||
```
|
||||
|
||||
Capped at **500 lines and 256KB, whichever binds first**, with `truncated` set
|
||||
so the UI can say so. Two caps because 500 lines of a container emitting 4KB
|
||||
JSON blobs is 2MB, and a line count alone does not stop it — the same reasoning
|
||||
that gave workflow logs both a per-line and a per-run cap.
|
||||
|
||||
Live following is deliberately absent. The browser console already offers a
|
||||
real terminal on the same server, where `docker logs -f` works properly with
|
||||
its own scrollback and cancellation. Building a second streaming path — a
|
||||
relay listener, proxy bus keys, a WebSocket upgrade and a cancellation story
|
||||
for a follow nobody closed — to duplicate that would be a large amount of
|
||||
machinery aimed at a capability already shipped. A bounded snapshot answers
|
||||
"why did this restart", which is the question that sends people to the console
|
||||
in the first place.
|
||||
|
||||
### Log reads are owner or admin only, and audited
|
||||
|
||||
Unlike workflow logs, these cannot be masked. A workflow's logs can be masked
|
||||
because the run injected the secrets and therefore knows their values. A
|
||||
container's stdout is arbitrary and may contain credentials nobody declared —
|
||||
a connection string in a startup banner, a token in a stack trace.
|
||||
|
||||
So log reads sit behind the same role check as control actions and are audited.
|
||||
A member who can see the fleet cannot read its logs. This is a deliberate
|
||||
access decision, not an oversight, and it is why log reading is not simply
|
||||
folded in with the read-only snapshot endpoints.
|
||||
|
||||
---
|
||||
|
||||
## REST API
|
||||
|
||||
```
|
||||
GET /api/servers/:id/workloads # stored snapshot
|
||||
POST /api/servers/:id/workloads/refresh # dispatch, then refetch
|
||||
POST /api/servers/:id/workloads/:wid/action # {"action":"start|stop|restart"} (owner|admin)
|
||||
GET /api/servers/:id/workloads/:wid/logs?tail= # (owner|admin)
|
||||
GET /api/workloads?image=&stack=&state= # fleet-wide
|
||||
```
|
||||
|
||||
`:wid` is a container ID or a unit name, URL-encoded. Unit names carry dots and
|
||||
`@`, which are legal in a path segment but not worth relying on unencoded.
|
||||
|
||||
`tail` is clamped to the 500-line cap server-side; a client asking for more
|
||||
gets 500, not an error.
|
||||
|
||||
---
|
||||
|
||||
## UI
|
||||
|
||||
Server detail gains a **Workloads** tab, ordered compose stacks first — grouped
|
||||
under the stack name — then loose containers, then units.
|
||||
|
||||
That ordering is not cosmetic. A stack is one thing to an operator even when it
|
||||
is six containers, and a flat list turns one decision into six rows. It is the
|
||||
same argument that groups the vulnerabilities board by CVE rather than by
|
||||
finding.
|
||||
|
||||
A `/workloads` fleet view answers "which servers run image X", which is the
|
||||
reason the snapshot is stored at all rather than fetched on demand and
|
||||
discarded.
|
||||
|
||||
Three rules that follow directly from the model:
|
||||
|
||||
- **Protected rows render their actions disabled, with the reason**, rather
|
||||
than offering a button whose refusal is already known.
|
||||
- **`DockerOK: false` reads "Docker not in use on this server"**, never an
|
||||
empty list, and `DockerError` when present is shown as a distinct problem.
|
||||
- State never reads by colour alone: every pill carries a distinct shape and a
|
||||
text label, matching the existing monitor and severity pills.
|
||||
|
||||
---
|
||||
|
||||
## Verification
|
||||
|
||||
**No automated tests.** The repository has none today, and by explicit
|
||||
instruction this feature adds none — no `*_test.go`, no frontend test files.
|
||||
A deliberate decision by the repository owner, recorded so the absence reads as
|
||||
a choice rather than an omission.
|
||||
|
||||
The behaviours that would otherwise have been tested are the ones that fail
|
||||
quietly, and the implementation plan carries a manual check for each:
|
||||
|
||||
- **Parser output against real command output.** `docker inspect` must yield
|
||||
`RestartCount`, health and the compose label as `Stack`; the `systemctl`
|
||||
exclusion filter must drop `systemd-*` and `user@*` while keeping
|
||||
`nginx.service`. Both are verified against a live host rather than a fixture.
|
||||
- **Protected-set computation.** `vantage-agent.service` marked,
|
||||
`nginx.service` not. Getting this wrong in the permissive direction lets a
|
||||
server stop its own agent, which is unrecoverable from the UI.
|
||||
- **Hash order-independence.** An ordering-sensitive hash resends the full list
|
||||
every 60 seconds, which is invisible except as traffic.
|
||||
- **Log capping in both directions.** 600 lines in → 500 out with `truncated`;
|
||||
a 300KB blob of fewer than 500 lines → capped, `truncated`. The second is the
|
||||
case a line-count-only implementation silently fails, and it fails by sending
|
||||
megabytes rather than by erroring.
|
||||
|
||||
---
|
||||
|
||||
## Failure modes
|
||||
|
||||
| Failure | Behaviour |
|
||||
| ------- | --------- |
|
||||
| Docker not installed | `DockerOK: false`, no error, UI reads "not in use" |
|
||||
| Docker installed, daemon down | `DockerOK: false` **plus** `DockerError` — different message, different fix |
|
||||
| Agent offline | 503 from the existing dispatcher. No queueing: a command whose owner died must fail loudly |
|
||||
| Action on a protected workload | Agent refuses; API answers 409 naming the reason |
|
||||
| `stop` exceeds its timeout | Real error surfaced, never a hopeful ack. Snapshot refreshed afterwards |
|
||||
| Container removed between snapshot and action | Docker's "No such container" surfaced and a refresh dispatched — this is what on-demand refresh is for |
|
||||
| Log exceeds either cap | Truncated, flagged, and stated in the UI |
|
||||
| Instance deleted | **`server_workloads` must be added to the control plane's instance-deletion collection list**, alongside sub-project A's two collections |
|
||||
|
||||
---
|
||||
|
||||
## Deliberately out of scope
|
||||
|
||||
- **Live log following.** The console already does it. See the logs section.
|
||||
- **Creating, deleting or updating containers and units.** This is a control
|
||||
and visibility surface, not a deployment tool — workflows already exist for
|
||||
changing what a server runs, with snapshots, audit and rollback.
|
||||
- **`docker exec` into a container.** The console reaches the host; exec from
|
||||
the control plane is a second remote-execution path with its own audit and
|
||||
authorisation story, and it belongs in its own spec if anywhere.
|
||||
- **Kubernetes and containerd.** The Docker collector shells to the `docker`
|
||||
CLI, so a node whose runtime is containerd or CRI-O reports nothing from it —
|
||||
`DockerOK: false`, correctly, since Docker genuinely is not in use. Covering
|
||||
those runtimes means a `crictl`/`nerdctl` collector, and talking to a
|
||||
Kubernetes API server is a different subsystem again. Neither is v1.
|
||||
- **Podman as a supported runtime.** Its `docker`-compatible CLI means an
|
||||
aliased install will largely work, and that is a happy accident rather than a
|
||||
claim: nothing here is tested against Podman and its `RestartCount` and
|
||||
compose-label behaviour are not verified.
|
||||
- **Windows.** No systemd, and a different container story.
|
||||
- **Image vulnerability scanning.** Sub-project C, which needs this spec's
|
||||
image list and sub-project A's findings model.
|
||||
@@ -1,297 +0,0 @@
|
||||
# API tokens and OpenAPI reference
|
||||
|
||||
Date: 2026-08-12
|
||||
Status: approved, ready for implementation planning
|
||||
|
||||
## Problem
|
||||
|
||||
The only programmatic credential the control plane issues is the ESO secrets-read
|
||||
bearer token, which reaches exactly one endpoint. Everything else requires a
|
||||
browser session cookie. There is therefore no supported way to drive Vantage from
|
||||
CI, a script, or infrastructure-as-code, and no machine-readable description of
|
||||
the REST API for anyone who wants to try.
|
||||
|
||||
This spec covers two deliverables that ship together: scoped API tokens, and an
|
||||
OpenAPI 3.1 document rendered as a live reference page. A Terraform provider is
|
||||
the intended follow-on and is explicitly out of scope here — it depends on both
|
||||
of these being settled, and it is a separate Go module with its own release
|
||||
cycle.
|
||||
|
||||
## Goals
|
||||
|
||||
- A person can mint a scoped, optionally expiring token and use it against the
|
||||
existing REST API with no new endpoints to learn.
|
||||
- A leaked token is bounded by role, by scope, and by expiry policy.
|
||||
- Offboarding a person removes their tokens as a side effect of removing them.
|
||||
- The API has a machine-readable description that cannot silently drift from the
|
||||
handlers it describes.
|
||||
- The reference page works on an air-gapped self-hosted install.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- Token editing. Role and scopes are immutable; rotation replaces amendment.
|
||||
- OAuth device flow or any browser-based authorisation grant.
|
||||
- Per-server or per-tag restrictions on a token.
|
||||
- Instance-owned service tokens that outlive their creator.
|
||||
- The Terraform provider.
|
||||
- General API rate limiting beyond the per-token limit described below.
|
||||
|
||||
## Part 1 — API tokens
|
||||
|
||||
### Token format and storage
|
||||
|
||||
A token is `vt_` followed by 32 random bytes, hex encoded. It is displayed once,
|
||||
at creation, and never again.
|
||||
|
||||
Only the SHA-256 hash is stored, in a unique index. This follows the precedent
|
||||
already set by `servers.agent_token_hash` and the ESO read token. bcrypt is
|
||||
deliberately not used: the value is full-entropy random rather than a
|
||||
user-chosen password, so a fast hash is sufficient, and a per-token salt would
|
||||
force a collection scan where an indexed lookup is wanted.
|
||||
|
||||
The first eight characters are stored in clear as `hint`, so the list can
|
||||
identify a token without revealing it.
|
||||
|
||||
### Authentication path
|
||||
|
||||
`auth.Middleware()` gains a fallback. When there is no `km_session` cookie it
|
||||
looks for `Authorization: Bearer vt_…`. Both paths end by placing a `*Session` in
|
||||
the gin context, so every handler, `auth.RequireRole`, `RequireActiveLicense`,
|
||||
`RequireFeature` and `actorFromCtx` continue to work unmodified.
|
||||
|
||||
```
|
||||
Session{
|
||||
UserID: token.UserID
|
||||
InstanceID: token.InstanceID
|
||||
Role: min(user.Role, token.Role) // owner > admin > member
|
||||
Email: user.Email
|
||||
TokenID: token.TokenID // "" for cookie sessions
|
||||
Scopes: token.Scopes // nil for cookie sessions
|
||||
}
|
||||
```
|
||||
|
||||
The effective role is recomputed on every request rather than frozen at
|
||||
creation. Demoting the user demotes the token with them. No caching is required
|
||||
because the user document is already read to confirm the user still exists.
|
||||
|
||||
The existing host guard applies identically. A token carries an `instance_id`,
|
||||
and a request arriving at a different instance's host is rejected exactly as a
|
||||
mismatched cookie session is. The tenant boundary must not have a token-shaped
|
||||
hole in it.
|
||||
|
||||
`last_used_at` is written best-effort and only when the stored value is more
|
||||
than 60 seconds old, so it does not become a Mongo write per request.
|
||||
|
||||
Rejections:
|
||||
|
||||
| Condition | Status | Body |
|
||||
| -------------------- | ------ | -------------------------------------- |
|
||||
| No credential at all | 401 | `not authenticated` |
|
||||
| Unknown token | 401 | `invalid token` |
|
||||
| Expired token | 401 | `code: token_expired` |
|
||||
| Owning user deleted | 401 | `invalid token` |
|
||||
| Missing scope | 403 | names the required scope |
|
||||
| Wrong instance host | 403 | `instance host mismatch` |
|
||||
|
||||
### Data model
|
||||
|
||||
New collection `api_tokens`, added to `services.ScopedCollections` so instance
|
||||
purge reaches it.
|
||||
|
||||
```
|
||||
instance_id string
|
||||
token_id string
|
||||
user_id string
|
||||
name string 1-64 chars, unique per user
|
||||
hint string first 8 chars of the plaintext
|
||||
token_hash string sha256
|
||||
role string owner|admin|member
|
||||
scopes []string
|
||||
expires_at *time.Time nil means never
|
||||
created_at time.Time
|
||||
last_used_at *time.Time
|
||||
created_by_ip string
|
||||
```
|
||||
|
||||
Indexes: unique on `token_hash`; compound on `(instance_id, user_id)`.
|
||||
|
||||
Deleting a user deletes their tokens as part of the same service call as
|
||||
`DeleteInstanceUser`, so offboarding is one action rather than two.
|
||||
|
||||
### Expiry policy
|
||||
|
||||
Expiry is optional by default: a token may be created with no expiry at all.
|
||||
Instance settings gain `api_token_max_days *int`, editable by owner and admin:
|
||||
|
||||
- `nil` — no cap; never-expire is allowed. This is the default, so an upgrade
|
||||
changes nothing.
|
||||
- `n > 0` — a new token must expire within `n` days, and a never-expire token is
|
||||
refused.
|
||||
|
||||
Changing the setting does not retroactively invalidate existing tokens; it is a
|
||||
policy on issuance. Tokens already outside the new cap are flagged in the UI so
|
||||
that someone can rotate them deliberately, rather than discovering the change
|
||||
when a pipeline breaks.
|
||||
|
||||
### Scopes
|
||||
|
||||
Eight resources, each with `:read` and `:write`. Write implies read on the same
|
||||
resource.
|
||||
|
||||
```
|
||||
servers keys secrets workflows
|
||||
monitors vulns workloads settings
|
||||
```
|
||||
|
||||
Scope enforcement is a single middleware, `RequireScopes()`, mounted once in the
|
||||
`/api` stack. It derives the required resource from the matched gin route
|
||||
pattern using a map, rather than from a per-route decorator: a route registered
|
||||
without a decorator would otherwise be unguarded, and this repo already prefers
|
||||
guards that come from where a route is mounted rather than from someone
|
||||
remembering.
|
||||
|
||||
- Cookie sessions skip the check entirely.
|
||||
- A token-authenticated request whose route pattern is absent from the map is
|
||||
denied with 403. Fail closed.
|
||||
- A startup check fails boot if any registered `/api` route pattern is missing
|
||||
from the map, so the failure surfaces at deploy rather than at the first call.
|
||||
|
||||
Deliberate placements:
|
||||
|
||||
- `keys:read` covers `GET /keys/:id/private-key`. Reading a private key is
|
||||
reading a key.
|
||||
- `secrets:read` does not cover `GET /api/secrets/:group/values`. That endpoint
|
||||
keeps its separate ESO bearer path and is unaffected by this work.
|
||||
- `workloads:write` covers both container control actions and log reads, which
|
||||
are already restricted to owner and admin.
|
||||
- The token endpoints themselves map to the `settings` resource: `GET
|
||||
/api/tokens` requires `settings:read`, and `POST` and `DELETE` require
|
||||
`settings:write`. A token can therefore mint or revoke tokens only when
|
||||
explicitly granted that scope, and never above its own role.
|
||||
|
||||
### Endpoints
|
||||
|
||||
```
|
||||
GET /api/tokens list; a member sees their own, owner|admin see all
|
||||
POST /api/tokens create; returns the plaintext once
|
||||
DELETE /api/tokens/:id revoke; own always, owner|admin any
|
||||
```
|
||||
|
||||
There is no `PUT`. Editing a token's role or scopes changes what a credential
|
||||
already deployed in a CI system can do, with no record of what it could do
|
||||
before. Rotation replaces amendment.
|
||||
|
||||
`POST` body: `name`, `role`, `scopes[]`, `expires_in_days` (omitted means never,
|
||||
and is refused when `api_token_max_days` is set).
|
||||
|
||||
Refusals: 400 for an unknown scope, 409 for a duplicate name for that user, 403
|
||||
for a role above the creator's own, 422 for an expiry beyond policy.
|
||||
|
||||
### Web UI
|
||||
|
||||
A new "API tokens" card in the Access group of `/settings`, alongside Members
|
||||
and single sign-on. Not a new nav entry — `/settings/instance` was folded back
|
||||
into `/settings` for precisely this reason, and the card lives in
|
||||
`web/components/settings/` with the others, reusing the shared `Field` and
|
||||
`inputClass`.
|
||||
|
||||
The card lists name, hint, role, scope chips, last used, and expiry with a
|
||||
distinct state for expired and for over-policy. Revoke is per row and confirms.
|
||||
|
||||
Create opens a modal. The plaintext is shown once in a `--well` block with
|
||||
copy-to-clipboard and an explicit line saying it will not be shown again.
|
||||
|
||||
Members see only their own rows. Owner and admin get an "All tokens" toggle.
|
||||
|
||||
`api_token_max_days` is a field on the same card, visible to owner and admin
|
||||
only.
|
||||
|
||||
### Audit
|
||||
|
||||
New events:
|
||||
|
||||
- `token.created`
|
||||
- `token.revoked`
|
||||
- `token.expired_use` — a rejected expired token, which is how a forgotten CI
|
||||
job becomes visible
|
||||
- `settings.token_policy_updated`
|
||||
|
||||
The actor is the human's email throughout, so `actorFromCtx` needs no change.
|
||||
Every existing audit event written during a token-authenticated request gains
|
||||
`via: "token:<name>"` in its detail, so the log distinguishes a person clicking
|
||||
from their credential acting.
|
||||
|
||||
### Rate limiting
|
||||
|
||||
Token-authenticated requests are limited per token in Redis at 600 per minute,
|
||||
answering 429 with `Retry-After`. Cookie sessions are untouched. This is narrow
|
||||
on purpose: it is not the general API rate-limiting project, only enough that a
|
||||
runaway script cannot take an instance down.
|
||||
|
||||
## Part 2 — OpenAPI and the reference page
|
||||
|
||||
### Generation
|
||||
|
||||
`swaggo/swag` v2, pinned, emitting OpenAPI 3.1. v1 emits Swagger 2.0, which
|
||||
Scalar renders poorly.
|
||||
|
||||
Handlers in `server/internal/api/*.go` gain annotation comments. Request and
|
||||
response bodies that are currently anonymous inline structs become named
|
||||
structs. This is real churn across roughly fifteen files and is the honest cost
|
||||
of choosing generation over a hand-written document.
|
||||
|
||||
The generated `server/internal/api/docs/openapi.json` is committed and embedded
|
||||
with `go:embed`, not generated during the image build: `server/Dockerfile`
|
||||
produces a `scratch` runtime from a Go build stage, and adding codegen there
|
||||
means putting the toolchain in the build image.
|
||||
|
||||
`server-deploy.yml` gains a check that regenerates the spec and runs
|
||||
`git diff --exit-code`. An annotation edited without regenerating fails the
|
||||
build. Without this check the annotations are worth less than a hand-written
|
||||
document, because they would drift while appearing authoritative.
|
||||
|
||||
### Serving
|
||||
|
||||
```
|
||||
GET /api/openapi.json the spec, session or token authenticated
|
||||
GET /api/docs HTML page loading a vendored Scalar bundle
|
||||
```
|
||||
|
||||
The Scalar standalone bundle is vendored under `server/internal/api/docs/`, with
|
||||
its version recorded in a comment beside it and refreshed by hand. No CDN:
|
||||
air-gapped self-hosted installs are supported, and a reference page that fails
|
||||
closed on an offline site is a support ticket.
|
||||
|
||||
Because the page is served by the instance itself, "Try it" acts against the
|
||||
reader's own API with their own session.
|
||||
|
||||
### Documented auth schemes
|
||||
|
||||
Three, kept distinct:
|
||||
|
||||
- `cookieAuth` — the `km_session` cookie.
|
||||
- `bearerAuth` — a `vt_…` API token.
|
||||
- The ESO secrets endpoint is marked as its own separate scheme, so nobody wires
|
||||
a personal access token into External Secrets Operator.
|
||||
|
||||
## Documentation
|
||||
|
||||
- `docsite/docs/reference/api-tokens.md`: creating a token, the scope table,
|
||||
curl examples, rotation, and the maximum-lifetime policy.
|
||||
- `CLAUDE.md`: the three token routes under REST API, the `api_tokens`
|
||||
collection, and a note that `openapi.json` is generated and CI-verified.
|
||||
|
||||
## Risks
|
||||
|
||||
- The anonymous-struct-to-named-struct conversion is the bulk of the work and
|
||||
touches handler code this feature otherwise has no business in.
|
||||
- The vendored Scalar bundle is a manual refresh that nobody will remember. The
|
||||
version comment is the only mitigation.
|
||||
- A scope map keyed on gin route patterns breaks if a route path is renamed. The
|
||||
boot-time completeness check is what turns that into a startup failure rather
|
||||
than a silent 403 in production.
|
||||
|
||||
## Follow-on work
|
||||
|
||||
A Terraform provider, as its own spec and plan, consuming the tokens and the
|
||||
OpenAPI document produced here.
|
||||
@@ -1,236 +0,0 @@
|
||||
# Instance rename in Vantage HQ
|
||||
|
||||
**Date:** 2026-08-12
|
||||
**Status:** approved, not yet implemented
|
||||
|
||||
## Problem
|
||||
|
||||
A cloud instance is named once, at creation, and never again. The name is
|
||||
chosen in the first thirty seconds of a customer's relationship with the
|
||||
product — before they have decided whether this is "Acme" or "Acme
|
||||
Production" — and it is the name that becomes their DNS host, appears in every
|
||||
sign-in link and heads every page of their control plane. Today the only way to
|
||||
change it is to create a second instance and move, or to open a support ticket
|
||||
that has no tooling behind it.
|
||||
|
||||
## What a rename is
|
||||
|
||||
One customer-initiated action on a **cloud** instance: a new name, from which a
|
||||
new slug is derived, which moves the instance to a new DNS host.
|
||||
|
||||
Name and slug move together. The slug is re-derived through
|
||||
`provision.BaseSlug`, so the rules that named the instance at creation are the
|
||||
rules that rename it — the same reserved-label list, the same 3–40 character
|
||||
bound, the same `Slugify` collapse of non-alphanumeric runs. There is no
|
||||
separate slug field for the customer to edit, because two fields invite the
|
||||
state where the name says one thing and the host says another, and that
|
||||
divergence is exactly what a rename exists to fix.
|
||||
|
||||
A licence binds an instance **UUID**, not a slug. A rename therefore issues no
|
||||
licence, calls Paddle not at all, and consumes no relink. This is the property
|
||||
that makes the whole feature cheap, and it should be stated in any future change
|
||||
that tempts someone to touch the licence from this path.
|
||||
|
||||
### What breaks, deliberately
|
||||
|
||||
- **The old host stops working.** The old slug is released the moment the rename
|
||||
commits; another account may take it. Bookmarks, saved sign-in links and any
|
||||
agent install one-liner that named the web host are stale. Agents themselves
|
||||
are unaffected — they dial `GRPC_HOST`, which is not per-tenant.
|
||||
- **The old host keeps working for up to 60 seconds.** `server/internal/auth/instancehost.go`
|
||||
caches slug-to-instance lookups for 60s, and admin has no path to invalidate
|
||||
another process's memory. The released slug can be claimed by another account
|
||||
inside that window, so for up to a minute a replica still maps that host to the
|
||||
previous tenant. No data is exposed — the host/session guard rejects a session
|
||||
belonging to a different instance — but the new owner's users can briefly reach
|
||||
the old tenant's instance on their own host, and see its login page rather than
|
||||
theirs. Adding a cross-service invalidation channel for a 60-second window is
|
||||
not worth the coupling.
|
||||
- **The customer must sign in again.** `km_session` is set with no `Domain`
|
||||
attribute, so it is host-only and does not follow the instance to its new
|
||||
subdomain. The UI says so rather than letting the customer discover it.
|
||||
|
||||
## Scope
|
||||
|
||||
| | Customer (owner or admin) | Staff |
|
||||
|---|---|---|
|
||||
| Cloud instance | rename, 24h cooldown | rename, no cooldown |
|
||||
| Self-hosted instance | refused, 400 | name only; there is no slug |
|
||||
| Cloud placeholder | refused, 409 | refused, 409 |
|
||||
|
||||
Self-hosted is refused on the customer side for the same reason the member
|
||||
endpoints refuse it: there is no control-plane row to write. The install is the
|
||||
customer's, on their own database, and admin cannot reach it. Staff may still
|
||||
correct the label on admin's own row, because that label is what staff search
|
||||
by.
|
||||
|
||||
## Data flow
|
||||
|
||||
Two writes, in this order:
|
||||
|
||||
1. **Control plane `instances`** — `{name, slug}`.
|
||||
2. **Admin `admin_instances`** — `{name, slug, renamed_at}`.
|
||||
|
||||
The control plane goes first because `instances.slug` carries the unique index,
|
||||
and that index is what actually decides a race between two accounts reaching for
|
||||
the same name. Deciding it anywhere else would be guessing.
|
||||
|
||||
If the second write fails, the first is rolled back best-effort — restoring the
|
||||
previous name and slug — and the request answers 500. Leaving them divergent
|
||||
would have HQ print a host that is not the host, which is worse than a failed
|
||||
rename.
|
||||
|
||||
## Backend
|
||||
|
||||
### `shared/provision/instance.go`
|
||||
|
||||
```go
|
||||
// ErrSlugTaken means the derived slug belongs to another instance.
|
||||
var ErrSlugTaken = errors.New("slug taken")
|
||||
|
||||
// RenameInstance changes an instance's name and re-derives its slug.
|
||||
func RenameInstance(ctx context.Context, db *mongo.Database, instanceID, name string) (*models.Instance, error)
|
||||
```
|
||||
|
||||
It lives beside `CreateInstanceWithID` so slug derivation keeps one home, and it
|
||||
behaves as that function's rules imply:
|
||||
|
||||
- `BaseSlug(name)` failures wrap `ErrNameRejected` — too short, too long,
|
||||
reserved.
|
||||
- The derived slug is compared against the instance's current one. If they are
|
||||
equal, only the name is written; a cosmetic capitalisation change is not a
|
||||
move, and must not fail on its own slug.
|
||||
- **No `-2` suffix loop.** Creation appends a counter because the customer is
|
||||
waiting on an instance and any free slug will do. A rename is a request for a
|
||||
specific host, and silently landing the customer on `acme-2` is a worse answer
|
||||
than refusing.
|
||||
- A duplicate-key error on the update surfaces as `ErrSlugTaken`, exactly as the
|
||||
create path treats it as "that slug is taken". The pre-check is a courtesy;
|
||||
the index is the boundary.
|
||||
|
||||
### `admin/internal/cloudprov`
|
||||
|
||||
```go
|
||||
func RenameInstance(ctx context.Context, instanceID, name string) (*sharedmodels.Instance, error)
|
||||
```
|
||||
|
||||
A thin wrapper over `provision.RenameInstance` on `db.ControlDB()`. It writes
|
||||
`instances` and nothing else, so admin's documented control-plane write boundary
|
||||
— `instances` and `users`, from `cloudprov` and `inject` only — is unchanged.
|
||||
|
||||
### `admin/internal/models`
|
||||
|
||||
`Instance` gains:
|
||||
|
||||
```go
|
||||
// RenamedAt is when this instance last changed name, and backs the 24h
|
||||
// customer cooldown. The cooldown is admin's policy, so it lives on admin's
|
||||
// row rather than in the control plane, which has no opinion about how often
|
||||
// a customer may move.
|
||||
RenamedAt *time.Time `bson:"renamed_at,omitempty" json:"renamed_at,omitempty"`
|
||||
```
|
||||
|
||||
A pointer because absent means "never renamed", and a zero `time.Time` would
|
||||
read as 1 January year 1 — far enough in the past that the cooldown is inert,
|
||||
but only by accident.
|
||||
|
||||
### `PUT /api/instances/:id/name` (customer)
|
||||
|
||||
Mounted in the `cust` group behind `auth.RequireAccountRole(owner, admin)`, and
|
||||
resolving the instance through `ownedInstance` like every other instance route,
|
||||
so another account's instance answers 404 rather than 403.
|
||||
|
||||
Body: `{"name": "..."}`, trimmed before use.
|
||||
|
||||
Refusals, in the order checked:
|
||||
|
||||
| Condition | Status | Body |
|
||||
|---|---|---|
|
||||
| `deployment != cloud` | 400 | `selfHostedRefusal`, the same constant and status the member endpoints already answer with |
|
||||
| `placeholder` | 409 | instance is not provisioned yet |
|
||||
| within 24h of `renamed_at` | 429 | includes the UTC time it unlocks |
|
||||
| `provision.ErrNameRejected` | 422 | the wrapped reason, verbatim |
|
||||
| `provision.ErrSlugTaken` | 409 | that name is already in use |
|
||||
|
||||
Success returns `{"instance_id", "name", "slug", "login_url"}` and writes an
|
||||
audit entry `instance.renamed` with detail `<old-slug> -> <new-slug>`, so the
|
||||
history of a host is answerable from the audit log alone.
|
||||
|
||||
`login_url` comes from the existing `loginURLFor(slug)`, which fills `{slug}`
|
||||
into `APP_LOGIN_URL` — the same builder the licence emails already use, rather
|
||||
than a second opinion about how a tenant host is spelled. It is empty when
|
||||
`APP_LOGIN_URL` is unset, and the portal then falls back to the host string it
|
||||
already composes from the slug in `InstanceRecord` and the instance page.
|
||||
|
||||
### `PUT /api/staff/instances/:id/name`
|
||||
|
||||
The same core, without the cooldown, actor recorded as the staff user. On a
|
||||
self-hosted instance it updates `admin_instances.name` only and does not call
|
||||
`cloudprov`.
|
||||
|
||||
## Frontend (`adminsite`)
|
||||
|
||||
### `lib/slug.ts`
|
||||
|
||||
A TypeScript mirror of `provision.Slugify` and the length/reserved checks, used
|
||||
only to preview the resulting host while the customer types. It carries the same
|
||||
warning as `web/lib/targets.ts`: it is a second implementation and must change in
|
||||
the same commit as the Go one. The preview can disagree with the server — the
|
||||
409 is the answer that counts.
|
||||
|
||||
### `components/RenamePanel.tsx`
|
||||
|
||||
An inline panel, not a modal — `adminsite` has no modal component, and the
|
||||
codebase's idiom for a destructive-ish action with one input is `RelinkPanel`:
|
||||
a control that expands in place inside a `Panel`.
|
||||
|
||||
Prefilled with the current name. Below the input, a live line reading
|
||||
`acme-ltd.vantage.hostxtra.co.uk` as the customer types, and a note that they
|
||||
will need to sign in again on the new host. Submit is disabled while the derived
|
||||
slug is unchanged or invalid.
|
||||
|
||||
It lives in an "Address" panel on `app/(customer)/instances/[id]/page.tsx`,
|
||||
rendered only when the instance is cloud and `account_role` is `owner` or
|
||||
`admin`. The staff instance page mounts the same component against the staff
|
||||
route.
|
||||
|
||||
`InstanceRecord` on the Overview page is not touched: it stays a summary, and
|
||||
the rename is a decision that deserves the detail page.
|
||||
|
||||
### After a successful rename
|
||||
|
||||
Invalidate `["account"]`, collapse the panel, and let the page redraw with the new
|
||||
name and host. The Console rail card shows the new host, with a note:
|
||||
|
||||
> This instance now lives at `acme-ltd.vantage.hostxtra.co.uk`. You will need to
|
||||
> sign in again there.
|
||||
|
||||
**No automatic redirect.** Sending the browser to the new host lands the customer
|
||||
on a login screen with no explanation, having just lost the HQ page they were
|
||||
standing on. The link is right there; they click it when they are ready.
|
||||
|
||||
## Verification
|
||||
|
||||
The repository has no Go test suite, so verification is build plus manual
|
||||
exercise, matching existing practice:
|
||||
|
||||
- `go build ./...` in `shared` and `admin`; `npm run build` in `adminsite`.
|
||||
- Rename a cloud instance; confirm `instances` and `admin_instances` agree on
|
||||
name and slug.
|
||||
- The new host serves a login page; the old host stops resolving to the instance
|
||||
within ~60 seconds.
|
||||
- A second rename within 24 hours answers 429.
|
||||
- A rename onto an occupied slug answers 409 and changes nothing.
|
||||
- A rename attempt on a self-hosted instance from the customer portal answers
|
||||
400, the same status and constant the member endpoints already answer with.
|
||||
- The audit log carries `instance.renamed` with both slugs.
|
||||
|
||||
## Out of scope
|
||||
|
||||
- Slug aliases or redirects from the old host. The control plane resolves one
|
||||
slug per instance, and an alias table is a second identity to keep correct for
|
||||
the sake of stale bookmarks.
|
||||
- Renaming from inside the control plane's own `/settings`. HQ owns instance
|
||||
identity, the same way it owns licences and `hq`-sourced users; a second
|
||||
writer would need the same collision handling and the same cooldown.
|
||||
- Any change to the licence, subscription or Paddle line items.
|
||||
Reference in New Issue
Block a user