docs: Cleanup old specs and plans

This commit is contained in:
2026-08-03 10:15:54 +01:00
parent 19ef773690
commit 1e2132c1a1
27 changed files with 0 additions and 35239 deletions
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -1,989 +0,0 @@
# Cloud Instance Creation — Phase 1: Identity
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Replace the control plane's global unique index on `users.email` with a per-instance one, and scope every lookup that relied on the global index, so one address can belong to several instances.
**Architecture:** The index change is safe only because the two unscoped `FindOne({email})` lookups are scoped in the same binary that performs the swap. The new compound index is created **before** the old one is dropped, so a failure at any point leaves a working constraint in place. The unscoped helper is deleted rather than left unused, and admin's one unscoped control-plane lookup — which has no instance to scope by — is removed entirely.
**Tech Stack:** Go 1.26, gin, MongoDB driver v2.8.0, `shared/indexes`, `shared/models`, `shared/provision`.
## Global Constraints
- **No automated Go tests.** Verification is by compiler, `grep`, and running built images against scratch databases. Every "confirm" step below is a command with expected output. This matches plans 0a through 4.
- **Never run `go` or `npm` on the host.** Everything runs in a container. The wrapper from earlier plans:
```sh
# /tmp/gorun.sh <module-dir> <command...>
DIR="$1"; shift
MSYS_NO_PATHCONV=1 docker run --rm -v "$(pwd)":/src -v vantage-gomod:/go/pkg/mod \
-v vantage-gocache:/root/.cache/go-build -w "/src/$DIR" \
golang:1.26 "$@"
```
- **`MSYS_NO_PATHCONV=1` on every `docker` call.** Git Bash rewrites container paths otherwise.
- **Run `go mod tidy` with `GOWORK=off`.** In workspace mode it drops `require` lines and the Docker build then fails with "missing go.sum entry".
- **`shared/` is consumed through `replace` directives** in `server`, `admin` and `sitesvc`. A change to `shared/` reaches all three on their next build; there is no version to bump.
- **All three service images must ship together.** An older image booting after this change would recreate `email_1`. `.gitea/workflows/server-deploy.yml` rebuilds every image on every push to `main`, so this is automatic — the hazard is only a partial manual rollout on the host.
- **This migration is one-way.** Once two users share an address across instances, `email_1` cannot be recreated. There is no rollback; fixes go forward.
- Nothing in this phase projects users, creates instances, or adds UI. Those are phases 2 and 3.
## Context this plan inherits
`CLAUDE.md` currently states that the unique index on user email is "a security property, not an optimisation", because `GetUserByEmail` does an unscoped `FindOne`. That statement is true today and stops being true in Task 1. Task 7 updates it in the same series of commits, and the replacement property is stronger: a scoped query cannot be ambiguous, whereas an index merely prevents the ambiguity from arising.
Spec: [`docs/superpowers/specs/2026-07-26-cloud-instance-creation-design.md`](../specs/2026-07-26-cloud-instance-creation-design.md), phase 1.
---
## File Structure
**Modified:**
| Path | Change |
| ----------------------------------- | -------------------------------------------------------------------------- |
| `shared/indexes/indexes.go` | compound `(instance_id, email)` unique index; idempotent drop of `email_1` |
| `shared/models/user.go` | `HQUserID` field, `AuthLocal`/`AuthOIDC`/`AuthHQ` constants |
| `server/internal/services/users.go` | `GetUserByEmail` deleted, `GetUserInInstanceByEmail` added |
| `server/internal/auth/local.go` | `resolveLoginInstance`, scoped sign-in |
| `server/internal/auth/oidc.go` | scoped lookup, cross-instance guard deleted |
| `admin/internal/auth/cloud.go` | **deleted** |
| `admin/internal/api/routes.go` | `/auth/login` points at `HandleCustomerLogin`; new staff route |
| `admin/internal/api/staff.go` | `staffCreateAccountUser` |
| `CLAUDE.md` | the index security-property paragraph, and the auth section |
**Created:** none.
---
### Task 1: Compound index and the drop
**Files:**
- Modify: `shared/indexes/indexes.go`
**Interfaces:**
- Consumes: nothing new.
- Produces: `indexes.EnsureCoreIndexes(ctx context.Context, db *mongo.Database) error` — unchanged signature, new behaviour. Called at boot by `server`, `sitesvc` and `admin`.
- [ ] **Step 1: Replace the body of `EnsureCoreIndexes` and add the drop helper**
Replace the whole file with:
```go
// Package indexes declares the MongoDB indexes more than one Vantage service
// depends on.
package indexes
import (
"context"
"errors"
"fmt"
"go.mongodb.org/mongo-driver/v2/bson"
"go.mongodb.org/mongo-driver/v2/mongo"
"go.mongodb.org/mongo-driver/v2/mongo/options"
)
// legacyUserEmailIndex is the global unique index on users.email that this
// package used to declare. It is dropped on sight.
const legacyUserEmailIndex = "email_1"
// indexNotFound is MongoDB's IndexNotFound error code. Two services booting at
// once can both decide to drop the legacy index; the loser must not treat that
// as a failure.
const indexNotFound = 27
// EnsureCoreIndexes declares the unique indexes on users and instances.
//
// users is unique on (instance_id, email), NOT on email alone. One address is
// one user WITHIN an instance; the same address may hold a user in several
// instances, because an account's people are projected into each instance they
// are granted access to.
//
// This is a security property, not an optimisation, and it is only sufficient
// because every lookup by email is scoped by instance. There is deliberately no
// unscoped lookup by email anywhere in the codebase: an unscoped FindOne would
// return an arbitrary one of several matching users, which on the login path
// means signing someone into a tenant that is not theirs. If you are about to
// add one, you are about to reintroduce that bug.
//
// Creating an index that already exists with the same specification is a no-op,
// so this is safe to call at every boot from every service.
func EnsureCoreIndexes(ctx context.Context, db *mongo.Database) error {
// Create the replacement BEFORE dropping the legacy index. A failure here
// leaves the old constraint in place, which is safe; a failure after the
// drop would leave the collection unconstrained, which is not.
if _, err := db.Collection("users").Indexes().CreateOne(ctx, mongo.IndexModel{
Keys: bson.D{{Key: "instance_id", Value: 1}, {Key: "email", Value: 1}},
Options: options.Index().SetUnique(true).SetName("instance_email_unique"),
}); err != nil {
return fmt.Errorf("users.instance_id+email index: %w", err)
}
if err := dropIndexIfExists(ctx, db.Collection("users"), legacyUserEmailIndex); err != nil {
return fmt.Errorf("drop users.%s: %w", legacyUserEmailIndex, err)
}
if _, err := db.Collection("instances").Indexes().CreateOne(ctx, mongo.IndexModel{
Keys: bson.D{{Key: "slug", Value: 1}},
Options: options.Index().SetUnique(true),
}); err != nil {
return fmt.Errorf("instances.slug index: %w", err)
}
return nil
}
// dropIndexIfExists drops name, treating "it was not there" as success whether
// that is discovered by listing or by racing another service to the drop.
func dropIndexIfExists(ctx context.Context, col *mongo.Collection, name string) error {
cur, err := col.Indexes().List(ctx)
if err != nil {
return err
}
var existing []struct {
Name string `bson:"name"`
}
if err := cur.All(ctx, &existing); err != nil {
return err
}
found := false
for _, i := range existing {
if i.Name == name {
found = true
break
}
}
if !found {
return nil
}
err = col.Indexes().DropOne(ctx, name)
if err == nil {
return nil
}
var srvErr mongo.ServerError
if errors.As(err, &srvErr) && srvErr.HasErrorCode(indexNotFound) {
return nil
}
return err
}
```
- [ ] **Step 2: Confirm it compiles**
Run:
```sh
sh /tmp/gorun.sh shared go build ./...
```
Expected: no output.
- [ ] **Step 3: Confirm the legacy index is not declared anywhere else**
Run:
```sh
grep -rn '"email"' --include=*.go shared/ server/ sitesvc/ admin/ | grep -i index
```
Expected: no matches. If sitesvc or the server declares its own `users.email` index, it would recreate what Task 1 drops.
- [ ] **Step 4: Commit**
```bash
git add shared/indexes/indexes.go
git commit -m "feat(shared): unique users index is (instance_id, email)
One address is one user within an instance, not globally, so an account's
people can be projected into every instance they are granted.
The replacement index is created before email_1 is dropped, so a failure
at any point leaves a working constraint. The drop is idempotent and
tolerates two services racing it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
```
---
### Task 2: `hq` fields on the user document
**Files:**
- Modify: `shared/models/user.go`
- Modify: `server/internal/models/user.go`
**Interfaces:**
- Consumes: nothing.
- Produces:
- `shared/models.AuthLocal = "local"`, `AuthOIDC = "oidc"`, `AuthHQ = "hq"`
- `shared/models.User.HQUserID string` — bson `hq_user_id,omitempty`
- the same three constants re-exported from `server/internal/models`, which is a thin alias file over `shared/models` and is what server code imports
Nothing writes `AuthHQ` or `HQUserID` in this phase. They land now so phases 2 and 3 do not have to change the shared module and rebuild every service again.
- [ ] **Step 1: Add the constants and the field**
In `shared/models/user.go`, after the `ValidRole` function, add:
```go
// Auth sources. A user's auth_source says who owns the row.
const (
AuthLocal = "local"
AuthOIDC = "oidc"
// AuthHQ marks a user projected from a Vantage HQ account. Its role,
// password and existence are owned by HQ, and the instance API refuses to
// change any of them locally — a role editable in two places is a role with
// two answers.
AuthHQ = "hq"
)
```
And in the `User` struct, add `HQUserID` immediately after `AuthSource`:
```go
type User struct {
ID bson.ObjectID `bson:"_id,omitempty" json:"_id,omitempty"`
UserID string `bson:"user_id" json:"user_id"`
InstanceID string `bson:"instance_id" json:"instance_id"`
Email string `bson:"email" json:"email"`
PasswordHash string `bson:"password_hash,omitempty" json:"-"`
Role string `bson:"role" json:"role"`
AuthSource string `bson:"auth_source" json:"auth_source"`
// HQUserID is the customer_users.user_id this row was projected from,
// absent on locally-created users.
HQUserID string `bson:"hq_user_id,omitempty" json:"hq_user_id,omitempty"`
CreatedAt time.Time `bson:"created_at" json:"created_at"`
LastLogin *time.Time `bson:"last_login,omitempty" json:"last_login,omitempty"`
}
```
- [ ] **Step 2: Re-export the constants from the server's alias file**
`server/internal/models/user.go` is a thin alias over `shared/models`, and server code imports that rather than the shared package directly. Add the auth sources alongside the roles it already re-exports:
```go
package models
import shared "gitea.hostxtra.co.uk/mrhid6/vantage/shared/models"
type User = shared.User
const (
RoleOwner = shared.RoleOwner
RoleAdmin = shared.RoleAdmin
RoleMember = shared.RoleMember
)
const (
AuthLocal = shared.AuthLocal
AuthOIDC = shared.AuthOIDC
AuthHQ = shared.AuthHQ
)
func ValidRole(role string) bool { return shared.ValidRole(role) }
```
- [ ] **Step 3: Confirm both compile**
Run:
```sh
sh /tmp/gorun.sh shared go build ./...
sh /tmp/gorun.sh server go build ./...
```
Expected: no output from either.
- [ ] **Step 4: Commit**
```bash
git add shared/models/user.go server/internal/models/user.go
git commit -m "feat(shared): auth_source constants and hq_user_id on User
Nothing writes them yet. They land now so phases 2 and 3 do not require a
second rebuild of every service that consumes the shared module.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
```
---
### Task 3: Scoped lookup in the user service
**Files:**
- Modify: `server/internal/services/users.go:65-75`
**Interfaces:**
- Consumes: `shared/indexes` from Task 1.
- Produces: `services.GetUserInInstanceByEmail(instanceID, email string) (*models.User, error)`.
- Removes: `services.GetUserByEmail`. Tasks 4 and 5 fix its two callers; the build will be red between this task and Task 5, which is expected and is why they are adjacent.
- [ ] **Step 1: Replace `GetUserByEmail`**
In `server/internal/services/users.go`, delete the whole `GetUserByEmail` function and put this in its place:
```go
// GetUserInInstanceByEmail finds a user by address WITHIN one instance.
//
// There is deliberately no unscoped lookup by email. users is unique on
// (instance_id, email), not on email alone, so an unscoped FindOne would return
// an arbitrary one of several matching users — which on the login path means
// signing someone into a tenant that is not theirs.
func GetUserInInstanceByEmail(instanceID, email string) (*models.User, error) {
email = strings.ToLower(strings.TrimSpace(email))
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
var u models.User
err := db.Col("users").FindOne(ctx, bson.M{
"instance_id": instanceID,
"email": email,
}).Decode(&u)
if err != nil {
return nil, err
}
return &u, nil
}
```
- [ ] **Step 2: Confirm the unscoped helper is gone and the build is red for the expected reason**
Run:
```sh
grep -rn "GetUserByEmail" --include=*.go .
```
Expected: exactly two matches, both call sites — `server/internal/auth/local.go` and `server/internal/auth/oidc.go`. No definition.
Run:
```sh
sh /tmp/gorun.sh server go build ./...
```
Expected: FAIL with `undefined: services.GetUserByEmail` at those two call sites. Any other error means something else was broken.
- [ ] **Step 3: Do not commit yet**
The build is red. Commit at the end of Task 5, when both callers are fixed. A commit that does not build is a commit nobody can bisect through.
---
### Task 4: Scoped local login
**Files:**
- Modify: `server/internal/auth/local.go:25-49`
**Interfaces:**
- Consumes: `services.GetUserInInstanceByEmail` from Task 3, `services.CountInstances` and `services.FirstInstance` from `server/internal/services/instances.go:57` and `:63`, `auth.InstanceFromHost` from `server/internal/auth/instancehost.go:53`.
- Produces: `resolveLoginInstance(c *gin.Context) (string, error)`, unexported, used only by this file.
**Behaviour change worth knowing:** signing in at the bare apex host stops working when more than one instance exists. Cloud sign-in is always on `<slug>.vantage.<tld>` — `APP_LOGIN_URL` fills `{slug}` in, so every link already points there — and self-hosted has exactly one instance, so both supported paths keep working. A bookmark to the apex login page on a multi-instance deployment will now get a 400 that names the cause.
- [ ] **Step 1: Add `resolveLoginInstance` and rewrite `HandleLocalLogin`**
In `server/internal/auth/local.go`, add `"fmt"` to the imports if it is not already there, then add above `HandleLocalLogin`:
```go
// resolveLoginInstance decides which instance a sign-in attempt belongs to.
//
// Cloud always answers from the host: every instance has its own subdomain, and
// APP_LOGIN_URL fills the slug in, so every sign-in link already points at one.
// Self-hosted has no subdomain and exactly one instance, because a licence
// binds one instance UUID.
//
// Anything else is refused rather than guessed. Picking an instance on someone's
// behalf is how you sign them into the wrong tenant.
func resolveLoginInstance(c *gin.Context) (string, error) {
if inst, ok := InstanceFromHost(c); ok {
return inst.InstanceID, nil
}
n, err := services.CountInstances()
if err != nil {
return "", err
}
if n != 1 {
return "", fmt.Errorf(
"cannot tell which instance this sign-in is for: %d instances exist and the host %q names none of them; sign in at your instance's own address",
n, c.Request.Host)
}
inst, err := services.FirstInstance()
if err != nil {
return "", err
}
return inst.InstanceID, nil
}
```
Then replace the body of `HandleLocalLogin` between the JSON bind and `SaveSession` with:
```go
instanceID, err := resolveLoginInstance(c)
if err != nil {
c.JSON(http.StatusBadRequest, gin.H{"error": err.Error()})
return
}
u, err := services.GetUserInInstanceByEmail(instanceID, body.Email)
if err != nil || !services.VerifyPassword(u, body.Password) {
c.JSON(http.StatusUnauthorized, gin.H{"error": "invalid credentials"})
return
}
```
The `SaveSession` call below it is unchanged: it already reads `u.InstanceID`.
- [ ] **Step 2: Confirm only the OIDC caller is left broken**
Run:
```sh
sh /tmp/gorun.sh server go build ./...
```
Expected: FAIL with `undefined: services.GetUserByEmail` at `internal/auth/oidc.go:130` only.
---
### Task 5: Scoped OIDC callback
**Files:**
- Modify: `server/internal/auth/oidc.go:129-141`
**Interfaces:**
- Consumes: `services.GetUserInInstanceByEmail` from Task 3.
- Produces: nothing new.
The cross-instance guard is deleted because it becomes unreachable: the lookup is now scoped to `instanceID`, so a user belonging to another instance is simply not found, and the OIDC callback provisions a new member — which is correct. OIDC is configured per instance, so only that instance's identity provider can reach this code with that instance's state.
- [ ] **Step 1: Replace the lookup and delete the guard**
In `server/internal/auth/oidc.go`, replace:
```go
email := strings.ToLower(claims.Email)
u, err := services.GetUserByEmail(email)
if err != nil {
u, err = services.CreateUser(instanceID, email, "", "member", "oidc")
if err != nil {
c.JSON(http.StatusInternalServerError, gin.H{"error": "provisioning failed"})
return
}
} else if u.InstanceID != instanceID {
c.JSON(http.StatusForbidden, gin.H{"error": "email belongs to a different organization"})
return
}
```
with:
```go
email := strings.ToLower(claims.Email)
// Scoped to the instance the callback state names, so an address that also
// exists in another instance is invisible here. That scoping replaces the
// cross-instance guard this code used to need: there is no longer a way for
// the lookup to return a user belonging to somebody else.
u, err := services.GetUserInInstanceByEmail(instanceID, email)
if err != nil {
u, err = services.CreateUser(instanceID, email, "", models.RoleMember, models.AuthOIDC)
if err != nil {
c.JSON(http.StatusInternalServerError, gin.H{"error": "provisioning failed"})
return
}
}
```
`services.CreateUser`'s signature is `CreateUser(instanceID, email, password, role, authSource string)` — the argument order above matches it, with the two string literals the old code passed replaced by the constants Task 2 added.
`oidc.go` already imports `gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/models`; confirm it before relying on the constants:
```sh
grep -n "server/internal/models" server/internal/auth/oidc.go
```
If that returns nothing, add the import rather than reverting to string literals — Task 2 exists so these two values have one spelling.
- [ ] **Step 2: Confirm the build is green**
Run:
```sh
sh /tmp/gorun.sh server go build ./...
```
Expected: no output.
- [ ] **Step 3: Confirm no unscoped email lookup survives anywhere in the server**
Run:
```sh
grep -rn "GetUserByEmail" --include=*.go .
```
Expected: no matches at all.
Run:
```sh
grep -rn 'FindOne(ctx, bson.M{"email"' --include=*.go server/
```
Expected: no matches.
**Coverage note.** The spec's phase-1 test 6 exercises this path end to end, which needs a working identity provider and is not reproducible in the container harness Task 7 uses. It is verified here by inspection and by the greps in Step 3 instead: the lookup is scoped by `instanceID`, which comes from `ConsumeStateInstance` and not from user input, and the deleted guard was the only other consumer of the unscoped helper. The first real OIDC sign-in after deployment is the confirming evidence — check that an existing SSO user still lands in their own instance before considering this closed.
- [ ] **Step 4: Commit Tasks 3, 4 and 5 together**
```bash
git add server/internal/services/users.go server/internal/auth/local.go server/internal/auth/oidc.go
git commit -m "feat(server): scope every user lookup by instance
users is unique on (instance_id, email) now, so an unscoped FindOne could
return an arbitrary one of several matching users. On the login path that
means signing someone into a tenant that is not theirs.
GetUserByEmail is deleted rather than left unused. Local sign-in resolves
its instance from the host, falling back to the single instance a
self-hosted deployment has, and refuses to guess otherwise. The OIDC
cross-instance guard goes: a scoped lookup cannot return another
instance's user, which is a stronger guarantee than the check it replaces.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
```
---
### Task 6: Remove admin's unscoped control-plane login
**Files:**
- Delete: `admin/internal/auth/cloud.go`
- Modify: `admin/internal/api/routes.go:30`, `admin/internal/api/routes.go:50-52`
- Modify: `admin/internal/api/staff.go`
**Interfaces:**
- Consumes: `auth.CreateCustomerUser(ctx, accountID, email, password string) error` from `admin/internal/auth/customer.go:32`.
- Produces: `POST /api/staff/accounts/:id/users`.
`HandleCloudLogin` authenticates against control-plane `users` with an unscoped `FindOne({email})`, and unlike the server's two lookups there is no instance in context to scope it by — HQ sign-in is not per-instance. It already falls through to `HandleCustomerLogin` whenever a `customer_users` row exists, which after phase 2 is every customer. Legacy cloud customers get an HQ login from staff, which is what the new endpoint is for; staff already attach those instances by hand per the spec README.
- [ ] **Step 1: Delete the file**
```sh
git rm admin/internal/auth/cloud.go
```
- [ ] **Step 2: Point `/auth/login` at the customer handler**
In `admin/internal/api/routes.go`, replace:
```go
r.POST("/auth/login", auth.HandleCloudLogin) // falls through to customer login
```
with:
```go
// Every customer authenticates against admin's own customer_users. There is
// deliberately no path that looks a customer up in the control plane by
// email alone: HQ sign-in names no instance, so such a lookup could not be
// scoped, and users.email is no longer globally unique.
r.POST("/auth/login", auth.HandleCustomerLogin)
```
- [ ] **Step 3: Add the staff route**
In `admin/internal/api/routes.go`, inside the `staff` group, immediately after the `staff.GET("/accounts/:id", staffGetAccount)` line, add:
```go
staff.POST("/accounts/:id/users", staffCreateAccountUser)
```
- [ ] **Step 4: Add the handler**
At the end of `admin/internal/api/staff.go`, add:
```go
// staffCreateAccountUser gives an account an HQ login.
//
// This is how a legacy cloud customer — one whose instance predates HQ accounts
// — gets into the portal, alongside the manual instance attach the spec README
// describes. It reuses CreateCustomerUser, so the row is unverified until the
// emailed link is opened and is rolled back if that email cannot be sent.
func staffCreateAccountUser(c *gin.Context) {
var body struct {
Email string `json:"email"`
Password string `json:"password"`
}
if err := c.ShouldBindJSON(&body); err != nil || body.Email == "" || len(body.Password) < 12 {
c.JSON(http.StatusBadRequest, gin.H{
"error": "email and a password of at least 12 characters are required"})
return
}
ctx := c.Request.Context()
accountID := c.Param("id")
if n, err := db.Admin("accounts").CountDocuments(ctx,
bson.M{"account_id": accountID}); err != nil || n == 0 {
c.JSON(http.StatusNotFound, gin.H{"error": "no such account"})
return
}
email := strings.ToLower(strings.TrimSpace(body.Email))
if err := auth.CreateCustomerUser(ctx, accountID, email, body.Password); err != nil {
c.JSON(http.StatusBadRequest, gin.H{"error": err.Error()})
return
}
s := auth.Current(c)
audit.Write(ctx, models.AuditEntry{
Actor: s.Email, Action: "customer_user.created", AccountID: accountID, Target: email})
c.JSON(http.StatusCreated, gin.H{"pending": true})
}
```
Confirm `strings` is imported in `staff.go`; add it if not:
```sh
grep -n '"strings"' admin/internal/api/staff.go
```
- [ ] **Step 5: Confirm the build is green and nothing still references the deleted handler**
Run:
```sh
grep -rn "HandleCloudLogin" --include=*.go .
```
Expected: no matches.
Run:
```sh
sh /tmp/gorun.sh admin go build ./...
```
Expected: no output. If `sharedmodels` is now an unused import in some file, remove that import line.
- [ ] **Step 6: Confirm admin has no unscoped control-plane user lookup left**
Run:
```sh
grep -rn 'db.Control("users")' --include=*.go admin/
```
Expected: no matches.
- [ ] **Step 7: Commit**
```bash
git add -A admin/
git commit -m "feat(admin): drop the unscoped control-plane login branch
HQ sign-in names no instance, so a lookup of control-plane users by email
alone cannot be scoped — and users.email is no longer globally unique, so
it would return an arbitrary match. Every customer authenticates against
customer_users instead.
Legacy cloud customers get an HQ login from staff via the new
POST /api/staff/accounts/:id/users, alongside the manual instance attach
the spec README already describes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
```
---
### Task 7: Documentation and end-to-end verification
**Files:**
- Modify: `CLAUDE.md`
**Interfaces:**
- Consumes: everything above.
- Produces: nothing.
This is the task that proves the change. With no test suite, this transcript is the only evidence, so run it in full rather than skimming it.
- [ ] **Step 1: Update `CLAUDE.md`**
In the **Auth and Orgs** section, replace the paragraph beginning "Unique indexes on user email and org slug are a **security property**" with:
```markdown
Unique indexes are a **security property**, not an optimisation. `users` is
unique on `(instance_id, email)` — one address is one user _within_ an instance,
and the same address may hold a user in several instances, because an account's
people are projected into each instance they are granted. This is sufficient only
because **every lookup by email is scoped by instance**; there is deliberately no
unscoped lookup anywhere, and adding one would let the login path return an
arbitrary one of several matching users. Instance slug, settings instance and ESO
token hash remain globally unique.
```
In the **Security** section, replace the "Unique indexes on user email, org slug…" bullet with:
```markdown
- Unique indexes on `(instance_id, email)`, instance slug, settings instance and the ESO token hash are load-bearing for tenant isolation. So is the absence of any unscoped lookup by email.
```
In the **MongoDB Collections** notes, add:
```markdown
- `users.auth_source` is `local`, `oidc` or `hq`. An `hq` user was projected from a Vantage HQ account and carries `hq_user_id`; HQ owns its role, password and existence.
```
- [ ] **Step 2: Build both images**
```sh
MSYS_NO_PATHCONV=1 docker build -q -f server/Dockerfile -t vantage-server:test .
MSYS_NO_PATHCONV=1 docker build -q -f admin/Dockerfile -t vantage-admin:test .
```
Expected: two image IDs. A "missing go.sum entry" failure here means `go mod tidy` was run in workspace mode.
- [ ] **Step 3: Start a scratch Mongo and Redis, and seed the OLD index**
Redis is not optional here: the server stores sessions in it, so every sign-in below fails without it.
```sh
MSYS_NO_PATHCONV=1 docker run -d --name vantage-idx-redis -p 6389:6379 redis:7
MSYS_NO_PATHCONV=1 docker run -d --name vantage-idx-mongo -p 27023:27017 mongo:7
MSYS_NO_PATHCONV=1 docker run --rm --add-host host.docker.internal:host-gateway mongo:7 \
mongosh "mongodb://host.docker.internal:27023/vantage_idx" --quiet --eval \
'db.users.createIndex({email:1},{unique:true}); db.getCollection("users").getIndexes().map(i=>i.name)'
```
Expected: output includes `email_1`. This reproduces a database that predates the change.
- [ ] **Step 4: Boot the server and confirm the swap**
```sh
MSYS_NO_PATHCONV=1 docker run -d --name vantage-idx-server -p 8091:8080 \
-e MONGO_URI=mongodb://host.docker.internal:27023 -e MONGO_DB=vantage_idx \
-e GRPC_HOST=localhost:9090 -e REDIS_ADDR=host.docker.internal:6389 \
--add-host host.docker.internal:host-gateway vantage-server:test
MSYS_NO_PATHCONV=1 docker run --rm --add-host host.docker.internal:host-gateway mongo:7 \
mongosh "mongodb://host.docker.internal:27023/vantage_idx" --quiet --eval \
'db.getCollection("users").getIndexes().map(i=>({name:i.name,key:i.key,unique:i.unique}))'
```
Expected: `instance_email_unique` present with key `{instance_id:1, email:1}` and `unique:true`; **no `email_1`**.
- [ ] **Step 5: Confirm a second boot is a no-op**
```sh
MSYS_NO_PATHCONV=1 docker restart vantage-idx-server
sleep 5
MSYS_NO_PATHCONV=1 docker logs vantage-idx-server 2>&1 | grep -i "index\|fatal" | tail -5
```
Expected: no index error and no fatal. The drop must tolerate the index already being gone.
- [ ] **Step 6: Bootstrap instance A and capture its user's password hash**
```sh
curl -s -X POST http://localhost:8091/auth/bootstrap \
-H 'Content-Type: application/json' \
-d '{"instance_name":"Alpha","email":"shared@example.com","password":"hunter2hunter2"}'
```
Expected: JSON with `instance_id` and `"slug":"alpha"`.
```sh
MSYS_NO_PATHCONV=1 docker run --rm --add-host host.docker.internal:host-gateway mongo:7 \
mongosh "mongodb://host.docker.internal:27023/vantage_idx" --quiet --eval \
'const u=db.users.findOne({email:"shared@example.com"}); print(u.user_id); print(u.password_hash)'
```
Expected: a UUID and a bcrypt hash. Keep both.
- [ ] **Step 7: Create instance B with the SAME address — the case that was impossible before**
```sh
MSYS_NO_PATHCONV=1 docker run --rm --add-host host.docker.internal:host-gateway mongo:7 \
mongosh "mongodb://host.docker.internal:27023/vantage_idx" --quiet --eval '
const a = db.users.findOne({email:"shared@example.com"});
const bId = UUID().toString().replace(/[{}]/g,"");
db.instances.insertOne({instance_id:bId, name:"Beta", slug:"beta", created_at:new Date()});
db.users.insertOne({
user_id: UUID().toString().replace(/[{}]/g,""),
instance_id: bId,
email: "shared@example.com",
password_hash: a.password_hash,
role: "owner",
auth_source: "local",
created_at: new Date()
});
print("beta instance " + bId);
print("users with that address: " + db.users.countDocuments({email:"shared@example.com"}));
'
```
Expected: `users with that address: 2`. Under the old global index this insert would have failed with E11000 — that failure is exactly what this phase removes.
- [ ] **Step 8: Confirm the compound index still refuses a duplicate WITHIN one instance**
```sh
MSYS_NO_PATHCONV=1 docker run --rm --add-host host.docker.internal:host-gateway mongo:7 \
mongosh "mongodb://host.docker.internal:27023/vantage_idx" --quiet --eval '
const a = db.users.findOne({email:"shared@example.com"});
try {
db.users.insertOne({user_id:"dup", instance_id:a.instance_id,
email:"shared@example.com", role:"member", auth_source:"local", created_at:new Date()});
print("FAIL: duplicate accepted");
} catch (e) { print("refused as expected: " + (e.code === 11000)); }
'
```
Expected: `refused as expected: true`. A `FAIL` line means the compound index is missing or not unique.
- [ ] **Step 9: Confirm each host signs in to its own instance — the whole point of the phase**
```sh
curl -s -X POST http://localhost:8091/auth/login -H 'Host: alpha.vantage.test' \
-H 'Content-Type: application/json' -c /tmp/alpha.jar \
-d '{"email":"shared@example.com","password":"hunter2hunter2"}'
curl -s http://localhost:8091/auth/me -H 'Host: alpha.vantage.test' -b /tmp/alpha.jar
```
Expected: `{"ok":true}`, then a body whose `instance` is **Alpha**.
```sh
curl -s -X POST http://localhost:8091/auth/login -H 'Host: beta.vantage.test' \
-H 'Content-Type: application/json' -c /tmp/beta.jar \
-d '{"email":"shared@example.com","password":"hunter2hunter2"}'
curl -s http://localhost:8091/auth/me -H 'Host: beta.vantage.test' -b /tmp/beta.jar
```
Expected: `{"ok":true}`, then a body whose `instance` is **Beta**, with a different `instance_id` from the Alpha response.
Two sign-ins, one address, one password, two different tenants. If both responses name the same instance, the lookup is not scoped.
- [ ] **Step 10: Confirm the apex host refuses rather than guesses**
```sh
curl -s -o /dev/null -w '%{http_code}\n' -X POST http://localhost:8091/auth/login \
-H 'Host: vantage.test' -H 'Content-Type: application/json' \
-d '{"email":"shared@example.com","password":"hunter2hunter2"}'
```
Expected: `400`. Then read the message:
```sh
curl -s -X POST http://localhost:8091/auth/login -H 'Host: vantage.test' \
-H 'Content-Type: application/json' \
-d '{"email":"shared@example.com","password":"hunter2hunter2"}'
```
Expected: an error naming both the instance count and the host. A `200` here would mean an arbitrary tenant was chosen.
- [ ] **Step 11: Confirm a wrong password still fails, on the right host**
```sh
curl -s -o /dev/null -w '%{http_code}\n' -X POST http://localhost:8091/auth/login \
-H 'Host: alpha.vantage.test' -H 'Content-Type: application/json' \
-d '{"email":"shared@example.com","password":"wrongwrongwrong"}'
```
Expected: `401`.
- [ ] **Step 12: Confirm a single-instance deployment still signs in on a bare host**
```sh
MSYS_NO_PATHCONV=1 docker run --rm --add-host host.docker.internal:host-gateway mongo:7 \
mongosh "mongodb://host.docker.internal:27023/vantage_idx" --quiet --eval \
'const b=db.instances.findOne({slug:"beta"}); db.users.deleteMany({instance_id:b.instance_id}); db.instances.deleteOne({slug:"beta"}); print(db.instances.countDocuments({}))'
```
Expected: `1`.
```sh
curl -s -X POST http://localhost:8091/auth/login -H 'Host: vantage.test' \
-H 'Content-Type: application/json' \
-d '{"email":"shared@example.com","password":"hunter2hunter2"}'
```
Expected: `{"ok":true}`. This is the self-hosted path, and it must keep working.
- [ ] **Step 13: Confirm admin boots and its login route still works**
```sh
MSYS_NO_PATHCONV=1 docker run -d --name vantage-idx-admin -p 8093:8083 \
-e ADMIN_MONGO_URI=mongodb://host.docker.internal:27023/vantage_idx_admin \
-e CONTROL_MONGO_URI=mongodb://host.docker.internal:27023/vantage_idx \
-e REDIS_ADDR=host.docker.internal:6389 \
-e LICENSE_SIGNING_KEY="$LICENSE_SIGNING_KEY" \
-e PUBLIC_URL=http://localhost:8093 -e ADMIN_ORIGIN=http://localhost:3004 \
--add-host host.docker.internal:host-gateway vantage-admin:test
sleep 5
curl -s http://localhost:8093/healthz
```
Expected: `{"ok":true}`. A boot failure here most likely means an unused-import error that `go build` caught but the image build did not, or a missing env var.
```sh
curl -s -o /dev/null -w '%{http_code}\n' -X POST http://localhost:8093/auth/login \
-H 'Content-Type: application/json' \
-d '{"email":"nobody@example.com","password":"hunter2hunter2"}'
```
Expected: `401`, not `500`. This proves `/auth/login` is wired to a live handler after `HandleCloudLogin` was deleted.
- [ ] **Step 14: Tear the scratch environment down**
```sh
MSYS_NO_PATHCONV=1 docker rm -f vantage-idx-server vantage-idx-admin vantage-idx-mongo vantage-idx-redis
```
- [ ] **Step 15: Commit**
```bash
git add CLAUDE.md
git commit -m "docs: users is unique per instance, not globally
The old index was load-bearing because two lookups were unscoped. Both
are scoped now and the unscoped helper is gone, so the property that
matters is the absence of any unscoped lookup by email. Says so, and
documents auth_source hq.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
```
---
## Done when
- `instance_email_unique` exists on `users`, `email_1` does not, and a second boot is a no-op.
- Two users share one address across two instances, and each signs in to their own.
- A duplicate address within one instance is still refused.
- The apex host refuses to guess when several instances exist, and still works when only one does.
- `grep -rn "GetUserByEmail"` and `grep -rn "HandleCloudLogin"` both return nothing.
- `admin` boots and `/auth/login` answers `401` rather than `500`.
- `CLAUDE.md` no longer claims `users.email` is globally unique.
**Not proven by this plan:** the OIDC sign-in path, which needs a real identity provider. Verify it manually on the first SSO sign-in after deployment — an existing SSO user must still land in their own instance.
## Not in this phase
`POST /api/instances`, the Free lifecycle, renewal, the notices, the reaper, the sitesvc cutover, account roles, invitations, instance membership, password propagation, and every UI change. Phases 2 and 3 get their own plans once this one lands.
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -1,928 +0,0 @@
# Control plane mobile responsiveness — Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Make `web/` (the Vantage control plane UI) usable on a phone — the sidebar becomes a hamburger-driven offcanvas below 1024px, tables become card stacks below 640px, and every fixed desktop layout collapses.
**Architecture:** A new client component `AppShell` owns the responsive chrome so `app/(app)/layout.tsx` stays a server component. `Sidebar.tsx` splits into a shared `SidebarContent` plus two containers (permanent aside, offcanvas drawer) so the nav exists in exactly one copy. The table card-stack lives in the `ui/Table.tsx` primitives via Tailwind `max-sm:` variants, so pages keep one markup tree and opt in with a `label` prop per cell.
**Tech Stack:** Next.js 16 (App Router), React 18, Tailwind 3.4, `clsx`. **No new dependencies.**
## Global Constraints
- **Scope is `web/` only.** Do not touch `site/`, `adminsite/`, `server/`, `admin/` or any Go code.
- **No hex colours anywhere.** Tailwind maps `var(--…)` tokens only. Use `bg-surface`, `border-border`, `text-text-secondary` etc. A literal `#` in a class is a defect. (`bg-black/60` is the one existing exception, already used by `Modal.tsx` for its backdrop — reuse it, do not introduce others.)
- **Breakpoints:** sidebar collapses below `lg` (1024px). Tables card-stack below `sm` (640px). Do not invent other breakpoints.
- **No new dependencies.** No headless-ui, no framer-motion.
- **Presentation only.** No API, route, query-key or data-shape changes.
- **Radius:** `rounded`, `rounded-lg`, `rounded-md` and `rounded-xl` all resolve to 46px via `tailwind.config.ts`. Prefer `rounded` in new code.
- Use `dvh`, not `vh`, for any new viewport-height value — mobile browser chrome makes `vh` overshoot.
- Indentation follows the file you are editing. `web/` is mixed: some files use 4 spaces (`Sidebar.tsx`, `keys/page.tsx`), others 2 (`servers/page.tsx`, `ui/*`). Match the file, do not reformat it.
- **There is no test framework in this repo.** No jest, no vitest, no playwright. Verification is `npx next lint`, `npx next build`, and targeted `grep` audits. Do not add a test framework.
- Run all commands from `d:\Development\Websites\vantage\web`.
---
### Task 1: Responsive table primitives
The card stack goes in the primitives, not the pages. Six pages render tables; giving each one a second markup tree would double the markup and drift on the first edit.
**Files:**
- Modify: `web/components/ui/Table.tsx` (whole file)
**Interfaces:**
- Consumes: nothing.
- Produces: `Td` gains an optional prop `label?: string`. Below `sm`, a `Td` with a `label` renders `<span>{label}</span>` before its children; a `Td` without one renders children alone, right-aligned. `Table`, `Thead`, `Tbody`, `Tr`, `Th` keep their existing signatures. Task 4 consumes `label`.
- [ ] **Step 1: Rewrite `web/components/ui/Table.tsx`**
Replace the entire file with:
```tsx
import { clsx } from "clsx";
import { HTMLAttributes, TdHTMLAttributes, ThHTMLAttributes } from "react";
/*
* Below sm the table stops being a table: the head is hidden, each row becomes
* a bordered card and each cell becomes a label/value pair. That lives here
* rather than in the six pages that render tables — two markup trees per page
* would drift apart on the first edit, and every one of those trees would mean
* the same thing.
*
* The mobile label uses Th's exact keyed-label idiom (mono, small, widely
* tracked, dimmed) because a key beside a value on a phone is the same device
* as a column head above it on a desktop.
*/
export function Table({ className, children, ...props }: HTMLAttributes<HTMLTableElement>) {
return (
<div className="overflow-x-auto">
<table
className={clsx("w-full border-collapse text-sm max-sm:block", className)}
{...props}
>
{children}
</table>
</div>
);
}
export function Thead({ className, children, ...props }: HTMLAttributes<HTMLTableSectionElement>) {
return (
<thead className={clsx("border-b border-border max-sm:hidden", className)} {...props}>
{children}
</thead>
);
}
export function Tbody({ className, children, ...props }: HTMLAttributes<HTMLTableSectionElement>) {
return (
<tbody
className={clsx(
"divide-y divide-border",
"max-sm:block max-sm:space-y-3 max-sm:divide-y-0 max-sm:p-3",
className
)}
{...props}
>
{children}
</tbody>
);
}
export function Tr({ className, children, ...props }: HTMLAttributes<HTMLTableRowElement>) {
return (
<tr
className={clsx(
"transition-colors hover:bg-surface-2/50",
"max-sm:block max-sm:rounded max-sm:border max-sm:border-border max-sm:bg-surface-2/40 max-sm:p-3",
className
)}
{...props}
>
{children}
</tr>
);
}
export function Th({ className, children, ...props }: ThHTMLAttributes<HTMLTableCellElement>) {
return (
<th
className={clsx(
// site/'s keyed-label idiom: mono, small, widely tracked, dimmed.
// A column head is a key, not prose.
// text-secondary, not tertiary: a column head is how you navigate the
// table, and tertiary lands under 4.5:1 at this size.
"px-4 py-3 text-left font-mono text-[0.68rem] uppercase tracking-[0.13em] text-text-secondary",
className
)}
{...props}
>
{children}
</th>
);
}
interface TdProps extends TdHTMLAttributes<HTMLTableCellElement> {
/**
* The column head this cell belongs to, shown beside the value below sm
* where the real head is hidden. Omit on a trailing action cell — an action
* needs no key, and the button then sits alone on its own row in the card.
*/
label?: string;
}
export function Td({ className, label, children, ...props }: TdProps) {
return (
<td
className={clsx(
"px-4 py-3 text-text-primary",
"max-sm:flex max-sm:items-start max-sm:gap-4 max-sm:px-0 max-sm:py-1.5",
// Exactly one justify class — clsx picks it. Emitting both and relying
// on string order would not work: Tailwind's output order decides which
// of two same-property utilities wins, not the order in this array.
label ? "max-sm:justify-between" : "max-sm:justify-end max-sm:pt-2.5",
className
)}
{...props}
>
{label && (
<span className="hidden font-mono text-[0.68rem] uppercase leading-5 tracking-[0.13em] text-text-secondary max-sm:inline">
{label}
</span>
)}
{children}
</td>
);
}
```
- [ ] **Step 2: Verify it compiles and lints**
```bash
npx tsc --noEmit
npx next lint
```
Expected: both clean. `tsc` may take ~30s. If `tsc --noEmit` errors on pre-existing issues unrelated to `Table.tsx`, note them and move on — only new errors matter.
- [ ] **Step 3: Commit**
```bash
git add web/components/ui/Table.tsx
git commit -m "feat(web): card-stack tables below sm"
```
---
### Task 2: Offcanvas sidebar
**Files:**
- Modify: `web/components/Sidebar.tsx` (whole file)
- Create: `web/components/AppShell.tsx`
- Modify: `web/app/(app)/layout.tsx` (whole file)
**Interfaces:**
- Consumes: `useAuth()` from `@/components/AuthProvider` returning `{ user, instance, isAdmin }`; `auth.logout()` from `@/lib/api`; `Logo` from `@/components/Logo`.
- Produces:
- `Sidebar.tsx` exports `SidebarContent({ onNavigate }: { onNavigate?: () => void })`, `Sidebar()` (permanent aside) and `SidebarDrawer({ open, onClose }: { open: boolean; onClose: () => void })`.
- `AppShell.tsx` exports `AppShell({ children }: { children: React.ReactNode })`.
- No later task depends on these names.
- [ ] **Step 1: Rewrite `web/components/Sidebar.tsx`**
Keep every icon component and the `navItems` array **exactly as they are** — do not retype the SVG path data, it is long and easy to corrupt. Change only from `export function Sidebar()` (line 135) to the end of the file, replacing it with the following. The file uses 4-space indentation.
```tsx
/** Shared by the permanent aside and the offcanvas drawer — one copy of the nav. */
export function SidebarContent({ onNavigate }: { onNavigate?: () => void }) {
const pathname = usePathname();
const { user, instance, isAdmin } = useAuth();
const visibleItems = navItems.filter((item) => !item.adminOnly || isAdmin);
const activeHref = visibleItems.reduce<string | null>((best, item) => {
const matches = pathname === item.href || pathname.startsWith(item.href + "/");
if (!matches) return best;
return best === null || item.href.length > best.length ? item.href : best;
}, null);
async function handleLogout() {
try {
await auth.logout();
} catch {}
window.location.href = "/login";
}
return (
<>
<div className="flex h-16 shrink-0 items-center gap-3 border-b border-border px-5">
<Logo className="h-8 w-8 text-logo" />
<div className="min-w-0">
<span className="block text-base font-extrabold leading-tight tracking-[-0.035em] text-text-primary">Vantage</span>
{instance && (
<span className="block truncate font-mono text-[0.68rem] uppercase tracking-[0.1em] text-text-secondary">{instance.name}</span>
)}
</div>
</div>
<nav className="flex-1 overflow-y-auto px-3 py-4">
<ul className="space-y-1">
{visibleItems.map((item) => {
const isActive = activeHref === item.href;
return (
<li key={item.href}>
<Link
href={item.href}
onClick={onNavigate}
// The active marker is an accent bar, the same device
// site/ uses to mark the chosen plan. A filled pill
// reads as a button you can press again.
className={clsx(
"relative flex items-center gap-3 rounded px-3 py-2.5 text-sm transition-colors",
isActive
? "bg-surface-2 font-semibold text-text-primary before:absolute before:inset-y-1 before:left-0 before:w-[2px] before:rounded-full before:bg-accent before:content-['']"
: "font-medium text-text-secondary hover:bg-surface-2 hover:text-text-primary",
)}
>
{item.icon}
{item.label}
</Link>
</li>
);
})}
</ul>
</nav>
<div className="shrink-0 border-t border-border px-4 py-3">
{user && (
<div className="mb-3">
<p className="truncate text-sm font-medium text-text-primary">{user.name || user.email}</p>
<p className="truncate text-xs text-text-secondary">
{user.email}
{user.role && <span className="ml-1 text-text-tertiary">· {user.role}</span>}
</p>
</div>
)}
<div className="flex items-center justify-between">
<p className="font-mono text-[0.68rem] uppercase tracking-[0.1em] text-text-secondary">Vantage v1.0</p>
{user && (
<button type="button" onClick={handleLogout} className="text-xs text-text-secondary transition-colors hover:text-danger">
Logout
</button>
)}
</div>
</div>
</>
);
}
/** The permanent sidebar. Below lg the drawer takes over. */
export function Sidebar() {
return (
<aside className="hidden h-screen w-60 shrink-0 flex-col border-r border-border bg-surface lg:flex">
<SidebarContent />
</aside>
);
}
/**
* The offcanvas below lg. Always mounted so the slide runs in both directions;
* closed it is inert (invisible + pointer-events-none) rather than unmounted.
*/
export function SidebarDrawer({ open, onClose }: { open: boolean; onClose: () => void }) {
const panelRef = useRef<HTMLDivElement>(null);
useEffect(() => {
if (!open) return;
const onKey = (e: KeyboardEvent) => {
if (e.key === "Escape") onClose();
};
window.addEventListener("keydown", onKey);
const previousOverflow = document.body.style.overflow;
document.body.style.overflow = "hidden";
panelRef.current?.focus();
return () => {
window.removeEventListener("keydown", onKey);
document.body.style.overflow = previousOverflow;
};
}, [open, onClose]);
return (
<div
className={clsx(
"fixed inset-0 z-50 lg:hidden",
open ? "visible" : "invisible pointer-events-none",
)}
>
<div
aria-hidden="true"
onClick={onClose}
className={clsx(
"absolute inset-0 bg-black/60 transition-opacity duration-200",
open ? "opacity-100" : "opacity-0",
)}
/>
<div
ref={panelRef}
id="app-sidebar-drawer"
role="dialog"
aria-modal="true"
aria-label="Navigation"
tabIndex={-1}
className={clsx(
"absolute inset-y-0 left-0 flex w-72 max-w-[85%] flex-col border-r border-border bg-surface outline-none transition-transform duration-200 ease-out",
open ? "translate-x-0" : "-translate-x-full",
)}
>
<SidebarContent onNavigate={onClose} />
</div>
</div>
);
}
```
Then update the import line at the top of the file (currently line 4) so `useEffect` and `useRef` are available:
```tsx
import { usePathname } from "next/navigation";
import { useEffect, useRef } from "react";
```
- [ ] **Step 2: Create `web/components/AppShell.tsx`**
```tsx
"use client";
import { useEffect, useRef, useState } from "react";
import { usePathname } from "next/navigation";
import { LicenseBanner } from "@/components/LicenseBanner";
import { Logo } from "@/components/Logo";
import { Sidebar, SidebarDrawer } from "@/components/Sidebar";
import { useAuth } from "@/components/AuthProvider";
function MenuIcon() {
return (
<svg className="h-6 w-6" fill="none" viewBox="0 0 24 24" stroke="currentColor" strokeWidth={1.5}>
<path strokeLinecap="round" strokeLinejoin="round" d="M3.75 6.75h16.5M3.75 12h16.5m-16.5 5.25h16.5" />
</svg>
);
}
/**
* Owns the responsive chrome so app/(app)/layout.tsx can stay a server
* component. Above lg this is the layout it always was; below lg the sidebar
* becomes an offcanvas behind the top bar's hamburger.
*/
export function AppShell({ children }: { children: React.ReactNode }) {
const [open, setOpen] = useState(false);
const pathname = usePathname();
const buttonRef = useRef<HTMLButtonElement>(null);
const { instance } = useAuth();
// A drawer that survives navigation would cover the page you just asked for.
useEffect(() => {
setOpen(false);
}, [pathname]);
function close() {
setOpen(false);
buttonRef.current?.focus();
}
return (
<div className="flex h-screen overflow-hidden">
<Sidebar />
<SidebarDrawer open={open} onClose={close} />
<div className="flex min-w-0 flex-1 flex-col overflow-y-auto">
<header className="sticky top-0 z-40 flex h-14 shrink-0 items-center gap-3 border-b border-border bg-surface px-3 lg:hidden">
<button
ref={buttonRef}
type="button"
onClick={() => setOpen(true)}
aria-label="Open navigation"
aria-expanded={open}
aria-controls="app-sidebar-drawer"
className="-ml-1 rounded p-2 text-text-secondary transition-colors hover:bg-surface-2 hover:text-text-primary"
>
<MenuIcon />
</button>
<Logo className="h-7 w-7 shrink-0 text-logo" />
<div className="min-w-0">
<span className="block text-sm font-extrabold leading-tight tracking-[-0.035em] text-text-primary">Vantage</span>
{instance && (
<span className="block truncate font-mono text-[0.62rem] uppercase tracking-[0.1em] text-text-secondary">{instance.name}</span>
)}
</div>
</header>
<main className="flex min-w-0 flex-1 flex-col">
<LicenseBanner />
{children}
</main>
</div>
</div>
);
}
```
- [ ] **Step 3: Rewrite `web/app/(app)/layout.tsx`**
```tsx
import { AuthProvider } from "@/components/AuthProvider";
import { AppShell } from "@/components/AppShell";
export default function AppLayout({
children,
}: {
children: React.ReactNode;
}) {
return (
<AuthProvider>
<AppShell>{children}</AppShell>
</AuthProvider>
);
}
```
`LicenseBanner` and `Sidebar` are no longer imported here — `AppShell` renders both.
- [ ] **Step 4: Verify**
```bash
npx tsc --noEmit
npx next lint
npx next build
```
Expected: all three succeed. `next build` is the one that matters — it catches a client component imported into a server component boundary.
- [ ] **Step 5: Sanity-check the scroll container**
Read `web/app/(app)/servers/[id]/console/page.tsx` around line 153 and 168. It uses `h-full`, which now resolves against `<main class="flex min-w-0 flex-1 flex-col">` rather than the old `<main class="flex-1 overflow-y-auto">`. Confirm the console page still has a height to fill; if `h-full` no longer resolves, change those two wrappers to `flex-1` instead. Task 7 revisits this file, so a note is acceptable here if you prefer to fix it there — but write the note down.
- [ ] **Step 6: Commit**
```bash
git add web/components/Sidebar.tsx web/components/AppShell.tsx "web/app/(app)/layout.tsx"
git commit -m "feat(web): offcanvas sidebar with hamburger below lg"
```
---
### Task 3: Page padding and header rows
**Files:**
- Modify: all 21 files under `web/app` and `web/components` containing `p-8`
- Modify: the title-plus-action header rows listed below
**Interfaces:**
- Consumes: nothing. Produces: nothing. Pure class edits.
- [ ] **Step 1: List every occurrence**
```bash
cd web && grep -rn "p-8" app components
```
Expected: 30 occurrences across 21 files.
- [ ] **Step 2: Replace each page-level `p-8` with `p-4 sm:p-6 lg:p-8`**
Apply to every occurrence **except** these two, which Task 6 and Task 7 handle and which need different values:
- `app/(app)/workflows/[id]/page.tsx:331` (the canvas `<main>`) — leave for Task 6.
- `app/(app)/servers/[id]/console/page.tsx:168` — leave for Task 7.
The inline loading states (`<div className="p-8 text-text-secondary">Loading…</div>`) get the same treatment: `className="p-4 text-text-secondary sm:p-6 lg:p-8"`.
Do this file by file with `Edit`. A blind `sed` would also hit `p-8` inside strings or unrelated contexts — check each match.
- [ ] **Step 3: Make title-plus-action header rows stack**
In each of these, change `className="mb-6 flex items-center justify-between"` to
`className="mb-6 flex flex-col gap-3 sm:flex-row sm:items-center sm:justify-between"`:
- `app/(app)/servers/page.tsx:75`
- `app/(app)/keys/page.tsx:112`
- `app/(app)/monitors/page.tsx:35`
- `app/(app)/workflows/page.tsx:32`
- `app/(app)/secrets/page.tsx:105`
- `app/(app)/secrets/[group]/page.tsx:251`
- `app/(app)/settings/notifications/page.tsx:153`
Leave `flex items-center justify-between` rows that are *inside* a card header or a table cell — those hold two small items and are fine at 390px. Only the page-top title/action rows change.
- [ ] **Step 4: Verify no unprefixed `p-8` survives**
```bash
cd web && grep -rn 'className="[^"]*\bp-8\b' app components | grep -v "sm:p-8\|lg:p-8"
```
Expected: exactly two lines — the two deferred to Tasks 6 and 7.
- [ ] **Step 5: Verify**
```bash
npx next lint && npx next build
```
Expected: both succeed.
- [ ] **Step 6: Commit**
```bash
git add web/app web/components
git commit -m "feat(web): responsive page padding and stacking page headers"
```
---
### Task 4: Label every table cell
**Files:**
- Modify: `web/app/(app)/servers/page.tsx:116-145`
- Modify: `web/app/(app)/keys/page.tsx:149-169`
- Modify: `web/app/(app)/monitors/page.tsx:79-98`
- Modify: `web/app/(app)/secrets/page.tsx:142-157`
- Modify: `web/app/(app)/secrets/[group]/page.tsx:141-150`
- Modify: `web/app/(app)/workflows/page.tsx:73-86`
- Modify: `web/app/(app)/workflows/[id]/runs/page.tsx:54-68`
- Modify: `web/app/(app)/audit/page.tsx:80-93`
- Modify: `web/app/(app)/keys/[id]/page.tsx:388-420`
- Modify: `web/app/(app)/servers/[id]/page.tsx:272-280` and `:588-605`
- Modify: `web/app/(app)/monitors/[id]/page.tsx:183-195`
- Modify: `web/components/settings/MembersCard.tsx:118-145`
**Interfaces:**
- Consumes: `Td`'s `label?: string` prop from Task 1.
- Produces: nothing.
- [ ] **Step 1: Add `label` to each `Td`, matching its `Th`**
For every table, the Nth `<Td>` in a `<Tr>` takes the text of the Nth `<Th>`. Where the `Th` is empty (`<Th />` — the trailing action column), the matching `Td` gets **no** `label`.
The mapping, `Th` order per file:
| File | Column labels, in order |
| --- | --- |
| `servers/page.tsx` | Hostname · IP Address · OS · Status · Last Seen · *(none)* |
| `keys/page.tsx` | Label · Fingerprint · Source · Assignments · Created · *(none)* |
| `monitors/page.tsx` | Name · Type · Target · Status · Latency · Last check |
| `secrets/page.tsx` | Group · Keys · Last Updated · *(none)* |
| `secrets/[group]/page.tsx` | Key · Value · Updated · *(none)* |
| `workflows/page.tsx` | Name · Targets · Steps · *(none)* |
| `workflows/[id]/runs/page.tsx` | Run · Status · Started · By · Servers |
| `audit/page.tsx` | Time · Event · Actor · Details |
| `keys/[id]/page.tsx` | Server · IP Address · Status · Assigned · Revoked · *(none)* |
| `servers/[id]/page.tsx` (updates table) | Package · Current · Available |
| `servers/[id]/page.tsx` (keys table) | Label · Fingerprint · Source · Status · Assigned · *(none)* |
| `monitors/[id]/page.tsx` | Started · Resolved · Cause |
| `MembersCard.tsx` | Email · Role · Sign-in · Last login · Actions |
Worked example — `servers/page.tsx` lines 116145 become:
```tsx
<Td label="Hostname">
<span className="font-medium text-text-primary">
{server.hostname}
</span>
</Td>
<Td label="IP Address">
<span className="font-mono text-text-secondary">
{server.ip_address}
</span>
</Td>
<Td label="OS">
<span className="text-text-secondary">{server.os_info}</span>
</Td>
<Td label="Status">
<StatusDot status={resolveStatus(server, latestVersion)} />
</Td>
<Td label="Last Seen">
<span className="text-text-secondary">
{server.last_seen
? formatLastSeen(server.last_seen)
: "Never"}
</span>
</Td>
<Td>
<Link href={`/servers/${server.server_id}`}>
<Button variant="ghost" size="sm">
View
</Button>
</Link>
</Td>
```
Note the last `Td` is unchanged — no `label`, so the "View →" button sits alone on its own row at the bottom of the card.
Second worked example — `MembersCard.tsx` line 142143, where `Td` already carries a `className`. Both props coexist:
```tsx
<Td label="Last login" className="text-text-secondary">{u.last_login ? new Date(u.last_login).toLocaleString() : "Never"}</Td>
<Td label="Actions" className="text-right">
```
`MembersCard`'s last column has a real `Th` ("Actions"), so unlike the others it **does** take a label.
- [ ] **Step 2: Verify no `Td` was missed**
```bash
cd web && grep -rn "<Td" app components | grep -v "label="
```
Expected: only the trailing action cells listed as *(none)* above — 7 of them (`servers`, `keys`, `secrets`, `secrets/[group]`, `workflows`, `keys/[id]`, `servers/[id]` keys table). Any other bare `<Td` is a miss.
- [ ] **Step 3: Verify**
```bash
npx next lint && npx next build
```
Expected: both succeed.
- [ ] **Step 4: Commit**
```bash
git add web/app web/components
git commit -m "feat(web): label table cells for the mobile card stack"
```
---
### Task 5: Modal bottom sheet and shared-component grids
**Files:**
- Modify: `web/components/ui/Modal.tsx:28-31`
- Modify: `web/components/monitors/MonitorForm.tsx:76,98,123,144`
- Modify: `web/components/workflows/StepPickerModal.tsx:132,168`
- Modify: `web/components/ui/Card.tsx:27`
**Interfaces:**
- Consumes: nothing. Produces: nothing.
- [ ] **Step 1: Make `Modal` a bottom sheet below `sm`**
In `web/components/ui/Modal.tsx`, replace lines 2834:
```tsx
<div className="fixed inset-0 z-50 flex items-end justify-center p-0 sm:items-center sm:p-4">
<div className="absolute inset-0 bg-black/60" onClick={onClose} />
<div
className={`relative z-10 w-full ${wide ? "sm:max-w-2xl" : "sm:max-w-md"} max-h-[85dvh] overflow-auto rounded rounded-b-none border border-b-0 border-border bg-surface shadow-panel sm:rounded sm:border-b`}
role="dialog"
aria-modal="true"
>
```
The `max-w-*` gains an `sm:` prefix so the sheet is full-width on a phone. `dvh` rather than `vh` because mobile browser chrome makes `vh` overshoot.
- [ ] **Step 2: Collapse the grids in `MonitorForm.tsx`**
- Line 76: `grid grid-cols-4 gap-2``grid grid-cols-2 gap-2 sm:grid-cols-4`
- Lines 98, 123, 144: `grid grid-cols-2 gap-4``grid grid-cols-1 gap-4 sm:grid-cols-2`
- [ ] **Step 3: Collapse the grids in `StepPickerModal.tsx`**
Lines 132 and 168: `grid grid-cols-2 gap-2.5``grid grid-cols-1 gap-2.5 sm:grid-cols-2`
- [ ] **Step 4: Let `CardHeader` wrap**
`web/components/ui/Card.tsx` line 27: `"mb-4 flex items-center justify-between"``"mb-4 flex flex-wrap items-center justify-between gap-2"`. Card headers hold a title and an action; at 390px they need to be allowed to wrap rather than crush the title.
- [ ] **Step 5: Verify**
```bash
npx next lint && npx next build
```
Expected: both succeed.
- [ ] **Step 6: Commit**
```bash
git add web/components/ui/Modal.tsx web/components/ui/Card.tsx web/components/monitors/MonitorForm.tsx web/components/workflows/StepPickerModal.tsx
git commit -m "feat(web): bottom-sheet modals and collapsing component grids"
```
---
### Task 6: Workflow builder
**Files:**
- Modify: `web/app/(app)/workflows/[id]/page.tsx:305-324` (header), `:329` (grid), `:331` (canvas), `:340` (column), `:372` (node), `:403` (inspector)
**Interfaces:**
- Consumes: nothing. Produces: nothing.
Below `lg` the fixed-height two-column grid is dropped entirely: single column, natural page flow. The `100dvh` arithmetic only makes sense at `lg`, where there is no mobile top bar above it.
- [ ] **Step 1: Let the header wrap (line 305)**
```tsx
<div className="flex flex-wrap items-center gap-3 border-b border-border bg-surface px-4 py-3">
```
and on line 312 change `className="ml-auto flex items-center gap-2"` to
`className="ml-auto flex flex-wrap items-center gap-2"`.
- [ ] **Step 2: Make the shell single-column below lg (line 329)**
```tsx
<div className="flex flex-1 flex-col lg:grid lg:h-[calc(100dvh-53px)] lg:grid-cols-[1fr_320px]">
```
`h-[calc(100vh-53px)]` becomes `lg:h-[calc(100dvh-53px)]``lg:` because the mobile top bar changes the arithmetic, and `dvh` because `vh` overshoots on mobile.
- [ ] **Step 3: Canvas padding (line 331)**
```tsx
<main className="overflow-auto bg-background bg-[radial-gradient(circle_at_1px_1px,theme(colors.border)_1px,transparent_0)] bg-[length:22px_22px] p-4 sm:p-6 lg:p-8">
```
- [ ] **Step 4: Let the node column and nodes be fluid (lines 340 and 372)**
Line 340:
```tsx
<div className="mx-auto flex w-full max-w-[340px] flex-col items-center">
```
Line 372 — the node itself. The wrapping `<div key={wfIdx} className="w-full">` on line 349 already constrains it, so the node just fills:
```tsx
className={`w-full cursor-pointer rounded border bg-surface p-3 ${isSelected ? "border-signal ring-2 ring-signal/40" : "border-border"}`}
```
- [ ] **Step 5: Turn the inspector into a bottom panel below lg (line 403)**
```tsx
<aside
className={`overflow-auto border-border bg-surface p-4 lg:block lg:border-l ${
selected === null || !selectedRef ? "hidden" : "block border-t max-lg:max-h-[60dvh]"
}`}
>
```
Below `lg` the inspector is hidden until a step is selected — an empty "Select a step to configure it" panel is noise on a phone — and when shown it sits under the canvas with a top border and a capped height. Above `lg` it is the left-bordered right rail it always was, always visible.
- [ ] **Step 6: Verify**
```bash
npx next lint && npx next build
```
Expected: both succeed.
- [ ] **Step 7: Commit**
```bash
git add "web/app/(app)/workflows/[id]/page.tsx"
git commit -m "feat(web): single-column workflow builder below lg"
```
---
### Task 7: Remaining fixed layouts
**Files:**
- Modify: `web/app/(app)/servers/[id]/page.tsx:164,495`
- Modify: `web/app/(app)/secrets/page.tsx:53`
- Modify: `web/app/(app)/workflows/[id]/runs/[runId]/page.tsx:~250`
- Modify: `web/app/(app)/servers/[id]/console/page.tsx:168` and its header rows
**Interfaces:**
- Consumes: nothing. Produces: nothing.
- [ ] **Step 1: `servers/[id]/page.tsx` line 164 — inventory grid**
`className="grid grid-cols-3 gap-2"``className="grid grid-cols-2 gap-2 sm:grid-cols-3"`
- [ ] **Step 2: `servers/[id]/page.tsx` line 495 — install one-liner**
`className="relative flex-1 min-w-64 rounded-lg border border-border bg-well px-4 py-2.5 font-mono text-sm"` → replace `min-w-64` with `min-w-0 overflow-x-auto`.
`min-w-64` is 256px of floor on a flex child; combined with a sibling copy button it pushes the row past a 390px viewport and scrolls the whole page sideways. `min-w-0` lets the box shrink and scroll its own content instead. Also check the parent flex row a few lines above and give it `flex-wrap` if the copy button ends up crushed.
- [ ] **Step 3: `secrets/page.tsx` line 53**
`className="grid grid-cols-2 gap-3"``className="grid grid-cols-1 gap-3 sm:grid-cols-2"`
- [ ] **Step 4: `workflows/[id]/runs/[runId]/page.tsx` — the step matrix**
Read the file around lines 240290. The matrix `<table>` has a `<th className="min-w-[240px] …">`. It is a genuine two-dimensional matrix (steps × servers) and must keep scrolling horizontally rather than stacking — stacking would destroy the information.
Confirm the `<table>` sits inside a wrapper with `overflow-x-auto`. If it does not, wrap it:
```tsx
<div className="overflow-x-auto">
<table >
</table>
</div>
```
If a wrapper already exists, leave it alone and note that in the commit body.
- [ ] **Step 5: `servers/[id]/console/page.tsx`**
Line 168: `className="flex h-full flex-col p-8"``className="flex h-full min-h-0 flex-1 flex-col p-4 sm:p-6 lg:p-8"`.
`flex-1` is added because Task 2 changed the parent `<main>` from `flex-1 overflow-y-auto` to `flex min-w-0 flex-1 flex-col`, so `h-full` alone may no longer resolve to anything. If Task 2 Step 5 recorded a note about this file, resolve it here.
Line 161's error state also has a bare `p-8` — Task 3 should already have handled it. Confirm it reads `p-4 sm:p-6 lg:p-8`.
Then read the connected-state toolbar below line 220 and add `flex-wrap` to any `flex items-center` row that holds three or more controls, so the console's chrome wraps instead of overflowing.
- [ ] **Step 6: Verify**
```bash
npx next lint && npx next build
```
Expected: both succeed.
- [ ] **Step 7: Commit**
```bash
git add web/app
git commit -m "feat(web): collapse remaining fixed layouts on small screens"
```
---
### Task 8: Final audit
**Files:** none modified unless the audit finds a miss.
- [ ] **Step 1: No unprefixed `p-8` remains**
```bash
cd web && grep -rn 'className="[^"]*\bp-8\b' app components | grep -v "sm:p-8\|lg:p-8"
```
Expected: no output.
- [ ] **Step 2: No unprefixed multi-column grid remains**
```bash
cd web && grep -rnoE '(class|className)="[^"]*(^|[" ])grid-cols-[2-9]' app components
```
Every hit must be a `grid-cols-2` that is genuinely fine at 390px (two short items side by side). Check each one and note the justification. Anything holding form controls or long text must gain a `grid-cols-1 sm:` prefix.
- [ ] **Step 3: No fixed pixel width escapes a breakpoint prefix**
```bash
cd web && grep -rnoE '(^|[" ])(w|min-w|max-w)-\[[0-9]{3,}px\]' app components
```
Expected: only `lg:`-prefixed hits, plus `max-w-[340px]` and `max-w-[1180px]` and `max-w-[300px]`, which are all *maximums* and shrink freely. A bare `w-[NNNpx]` or `min-w-[NNNpx]` without a prefix is a defect — except `min-w-[240px]` in the run-detail matrix, which is deliberate (Task 7 Step 4).
- [ ] **Step 4: No hex colours were introduced**
```bash
cd web && git diff main --stat && git diff main -- app components | grep -nE '^\+.*#[0-9a-fA-F]{3,8}\b'
```
Expected: no output from the grep. Tailwind in this app maps `var(--…)` tokens only.
- [ ] **Step 5: Full build and lint**
```bash
npx next lint
npx next build
```
Expected: both succeed with no new warnings.
- [ ] **Step 6: Read the diff end to end**
```bash
git diff main -- web/
```
Check for: an accidentally deleted SVG path, a `Td` whose `label` does not match its `Th`, indentation reformatted in a file that used the other convention.
- [ ] **Step 7: Commit any fixes**
```bash
git add web
git commit -m "fix(web): mobile audit corrections"
```
If the audit found nothing, skip this step — do not create an empty commit.
---
## Self-review notes
**Spec coverage:** shell → Task 2; tables → Tasks 1 and 4; padding and headers → Task 3; modal → Task 5; workflow builder → Task 6; remaining fixed layouts → Task 7; verification → Task 8 plus a verify step in every task.
**Known limitation:** there is no test framework and no running backend in this environment, so no task can prove a page *looks* right — only that it compiles, lints, and contains no pattern known to break at 390px. The first person to open this on a phone should expect to find something. That is a property of the verification approach chosen in the spec, not a gap in the plan.
File diff suppressed because it is too large Load Diff
@@ -1,410 +0,0 @@
# Spec 3 — Admin Backend
Date: 2026-07-24
Status: Design approved, not implemented
Depends on: spec 0a, spec 0b, spec 1 (`licensing-core`)
Ships: with spec 4 (the site it serves). Blocks specs 4 and 5.
## Context
A fourth Go service, `admin/`, owning customers, instances, licences and
subscriptions. It is the only service that holds the signing key.
Its data model separates two things the control plane deliberately does not know
about:
- **Account** — a paying customer. Holds a Paddle customer, a billing email, and
one or more instances.
- **Instance** — one deployment. Cloud instances mirror a control-plane
`Instance` row; self-hosted instances exist only here, because the customer's
database is theirs and we cannot see it.
## Goals
1. Issue, store and re-issue licences, with full history.
2. Inject licences into cloud instances.
3. Serve both staff and customers, with the right things hidden from each.
4. Never be a runtime dependency of a Vantage instance. If admin is down,
every instance keeps working; only purchasing and renewals stop.
## Non-goals
- The UI. Spec 4.
- Paddle. Spec 5. This spec defines the `subscriptions` table and the issuance
functions that spec 5's webhooks call, and nothing more.
- Rebuilding billing management. Card changes, invoices and cancellation go to
Paddle's own customer portal.
## Design
### Module
```
admin/
├── go.mod # replace => ../shared
├── cmd/main.go
└── internal/
├── api/ # gin handlers
├── auth/ # staff, cloud-customer and local-customer sessions
├── db/ # two connections: admin DB and control-plane DB
├── models/ # admin-owned documents
├── licensing/ # issuance, renewal, relink
├── inject/ # control-plane writes
├── mail/ # licence delivery
└── paddle/ # spec 5 lands here
```
Port `8083`. In `deploy/docker-compose.site.yml` only — like sitesvc, admin is
**excluded from the self-hosted deployment**. A self-hosted customer runs
instances, not the licensing authority.
### Two database connections
The service holds two:
- `ADMIN_MONGO_URI` — its own database, `vantage_admin`. Sole owner.
- `CONTROL_MONGO_URI` — the control plane's database, used to write licence
fields onto cloud instance documents and to authenticate cloud customers.
The control-plane connection uses the `Instance` and `User` structs from
`shared` (spec 0a). This is what makes direct writes safe: there is no admin-side
copy of the document shape to drift, which is the coupling hazard sitesvc used to
carry.
Admin's control-plane access is **narrow by construction**: it reads
`instances` and `users`, and it writes exactly three fields on `instances`. The
Mongo credential it is given should be scoped to that where the deployment allows
it. It must never write to any other collection.
### Data model
```go
type Account struct {
ID bson.ObjectID
AccountID string // uuid
Name string
BillingEmail string
PaddleCustomerID string // empty until first checkout
Status string // active | suspended
CreatedAt time.Time
}
type Instance struct {
ID bson.ObjectID
InstanceID string // for cloud: equals the control-plane instance_id
// for self-hosted: the UUID the customer pasted
AccountID string
Name string
Slug string // cloud only; the subdomain label
Deployment string // cloud | self_hosted
Tier string
Status string // awaiting_link | active | lapsed | cancelled
CurrentLicense string // licence ID
RelinkCount int // reset each term
CreatedAt time.Time
}
type License struct {
ID bson.ObjectID
LicenseID string
InstanceID string
AccountID string
Tier string
Deployment string
Limits license.Limits // snapshot
Features []string // snapshot
IssuedAt time.Time
ExpiresAt time.Time
Blob string
SupersededBy string // licence ID, when replaced
IssuedBy string // staff user, "system", or "paddle:<event id>"
Reason string // new | renewal | tier_change | relink | manual
}
type Subscription struct {
ID bson.ObjectID
SubscriptionID string
AccountID string
InstanceID string
PaddleSubscriptionID string
PaddlePriceID string
Tier string
Term string // monthly | annual
Status string // active | past_due | cancelled | awaiting_link
CurrentPeriodEnd time.Time
}
type Plan struct {
Tier string
Name string
Deployment string
Limits license.Limits
Features []string
PaddleProductID string
PaddlePriceIDs map[string]string // "monthly" | "annual"
Active bool
}
```
Collections: `accounts`, `admin_instances`, `licenses`, `subscriptions`,
`plans`, `staff_users`, `customer_users`, `admin_audit`.
Unique indexes: `accounts.account_id`, `admin_instances.instance_id`,
`licenses.license_id`, `subscriptions.paddle_subscription_id`, `plans.tier`,
`staff_users.email`, `customer_users.email`.
`admin_instances.instance_id` unique is load-bearing: it is what stops the same
self-hosted UUID being linked to two accounts.
**Licences are append-only.** A renewal writes a new row and sets
`SupersededBy` on the old one. Nothing is ever edited or deleted. When a support
question arrives about why a customer's instance stopped working on a given
date, the answer is in the table.
`plans` holds tier contents so they change without a deploy, seeded from the
table in spec 1. Every issued licence snapshots the plan, so editing a plan never
changes an existing licence — the same rule as `workflow_runs.steps_snapshot`.
### Issuance
```go
func Issue(ctx, instanceID, tier, term, reason, issuedBy string) (*models.License, error)
```
1. Load the instance and its account.
2. Load the plan for `tier`; refuse if `plan.Deployment != instance.Deployment`.
**This is the check that makes Free cloud-only** — Free's plan is
`deployment: cloud`, so it can never be issued to a self-hosted instance.
3. Build the payload with the instance's UUID bound in, `ExpiresAt` from the term
plus a **3-day grace** so a renewal webhook arriving slightly late does not
create a gap.
4. Sign with `LICENSE_SIGNING_KEY`.
5. Insert the licence row; set `SupersededBy` on the previous one; update
`instance.CurrentLicense` and `instance.Tier`.
6. If cloud, inject. If self-hosted, email the blob and make it downloadable.
7. Write an `admin_audit` entry.
Steps 5 and 6 are not transactional. Order matters: **record first, deliver
second.** A licence recorded but not delivered is recoverable — the customer
downloads it. A licence delivered but not recorded is a support mystery.
### Free tier rule
One Free instance per account, enforced in `Issue`: refuse a second Free instance
for an account that already has one that is not `cancelled`. Additional
instances must be paid.
### Injection
```go
func InjectCloud(ctx, instanceID string, lic *models.License) error
```
Writes `license_blob`, `license_tier`, `license_expiry` onto the control-plane
`instances` document via a single `UpdateOne`. Idempotent, retryable, and safe to
re-run.
Retries three times with backoff; on final failure the licence stays recorded and
`instance.Status` is set to `active` regardless, with the failure logged and
surfaced as a staff alert. A **reconciliation job runs every 15 minutes**,
comparing each cloud instance's `CurrentLicense` against the blob actually stored
in the control plane, and re-injecting on mismatch. That job, not the webhook, is
what guarantees eventual consistency.
The control-plane instance caches licence state for 60 seconds (spec 2), so an
injection takes effect within a minute without a restart.
### Self-hosted linking
The flow, end to end:
```
Customer runs /setup on their own install → instance UUID generated and shown
Customer buys Self Hosted in the admin site → subscription created,
status awaiting_link
Customer pastes the UUID into the admin site → admin_instances row created,
status active
Admin issues the licence with that UUID bound in
Customer downloads the .lic file or copies the blob
Customer pastes it into /settings/license on their install
```
Validation on link: the UUID must parse as a UUID, must not already exist in
`admin_instances`, and must not collide with a cloud instance ID. A duplicate
returns "That instance ID is already linked to an account" without revealing
which — it is a small enumeration surface but there is no reason to leave it
open.
### Relink
A rebuilt server has a new UUID. `POST /api/instances/:id/relink` with the new
UUID:
- Allowed **3 times per term**, `RelinkCount` reset on renewal.
- Updates `admin_instances.instance_id`, issues a replacement licence for the
**remaining term** with `reason: relink`, supersedes the old one.
- The old licence is not revoked — it cannot be, offline verification has no
revocation. It simply no longer matches any UUID the customer controls, and its
binding stops it being useful on a different machine anyway.
- Beyond 3, the endpoint returns a message directing the customer to support, and
staff can relink without limit.
`RelinkCount` is the abuse signal, not the abuse prevention. Its real job is to
put a human in front of the fourth attempt.
### Authentication
Three identities, three paths, one session store (Redis, `admin_session`
cookie, 24h).
**Staff**`staff_users`, local email plus bcrypt. Full access. Created by CLI
only; there is no staff signup.
**Cloud customers** — authenticate against the control plane's `users`
collection with the credentials they already use. Admin looks the user up
by email, checks bcrypt, resolves their control-plane instance, then resolves the
account that owns it.
Two consequences, stated plainly because they are real:
1. A cloud user's control-plane password now also unlocks billing. Any password
change or compromise has a wider blast radius than before.
2. Only users with control-plane role `owner` may sign in to the admin site.
`admin` and `member` are refused. Billing is an owner concern.
Mitigations: rate-limit to 5 attempts per email per 15 minutes and 20 per IP per
hour; log every attempt to `admin_audit`; return an identical error for unknown
email and wrong password.
**Self-hosted customers**`customer_users`, local email plus bcrypt at cost 12,
created during purchase, scoped to one account. Email verification reuses the
pattern sitesvc already proved: 32 random bytes, only the SHA-256 hash stored,
24-hour expiry, TTL index.
A single email address could in principle be both a cloud user and a
self-hosted customer user. `customer_users` is checked first; if it matches, that
identity wins. Documented so the behaviour is chosen rather than emergent.
### API
Staff:
```
GET /api/staff/accounts list, search
POST /api/staff/accounts
GET /api/staff/accounts/:id
GET /api/staff/instances filter by account, deployment, status, expiry
POST /api/staff/instances/:id/issue manual issue or reissue
POST /api/staff/instances/:id/relink no limit
GET /api/staff/licenses full history, filterable
GET /api/staff/plans
PUT /api/staff/plans/:tier
GET /api/staff/audit
GET /api/staff/health/injection reconciliation status and failures
```
Customer:
```
GET /api/account own account and instances
POST /api/instances/link self-hosted UUID link
POST /api/instances/:id/relink rate-limited
GET /api/instances/:id/license current licence metadata
GET /api/instances/:id/license/download .lic file
GET /api/subscriptions status, next renewal
POST /api/billing/portal Paddle portal redirect (spec 5)
```
Every customer handler resolves the account from the session and scopes by it.
The scoping is enforced by a helper every handler calls, not by each handler
remembering — the same deny-by-default reasoning as spec 2's middleware.
### Configuration
| Variable | Required | Notes |
|---|---|---|
| `ADMIN_MONGO_URI` | yes | admin's own database; name read from the URI path, refused if absent |
| `CONTROL_MONGO_URI` | yes | control-plane database, for injection and cloud auth |
| `REDIS_ADDR` | yes | sessions |
| `LICENSE_SIGNING_KEY` | yes | ECDSA P-384 private key, base32 (lk PrivateKey.ToB32String). **Boot fails without it** — a licensing service that cannot sign is worse than one that is down, because it looks healthy |
| `PUBLIC_URL` | yes | for verification and licence links |
| `SMTP_*` | yes | licence delivery |
| `ADMIN_ORIGIN` | yes | CORS allow-list |
| `TRUST_PROXY` | no | only behind a proxy that overwrites `X-Forwarded-For` |
| Paddle variables | spec 5 | |
### Backfill
Licences issued by `lkctl` during the spec 12 period exist only as blobs.
A one-shot `admin backfill --from=blobs.json` parses each with
`license.Parse`, creates the account, instance and licence rows, and marks them
`reason: manual`. Run once when admin goes live.
## Testing
**Issuance:**
1. `Issue` produces a licence that `license.Verify` accepts for that instance.
2. Deployment mismatch (Free plan, self-hosted instance) is refused.
3. A second Free instance for the same account is refused; a third paid one is
allowed.
4. Renewal supersedes the previous licence and leaves it in the table.
5. The issued licence snapshots the plan; editing the plan afterwards does not
change the issued licence.
6. Grace period: `ExpiresAt` is term end plus 3 days.
**Injection:**
7. `InjectCloud` writes all three fields; the control plane then reports `valid`.
8. Injection is idempotent across two calls.
9. Injection failure leaves the licence recorded and flags the instance.
10. The reconciliation job detects a control-plane blob that does not match
`CurrentLicense` and re-injects.
**Linking and relink:**
11. Linking an unknown UUID succeeds; linking one already linked is refused.
12. Relink issues a licence for the *remaining* term, not a fresh full term.
13. The fourth relink in a term is refused for a customer and allowed for staff.
14. `RelinkCount` resets on renewal.
**Auth:**
15. Cloud owner signs in with control-plane credentials; `admin` and `member`
roles are refused.
16. Unknown email and wrong password return identical errors and timing is not a
meaningful oracle.
17. Rate limits trigger at the documented thresholds.
18. Self-hosted customer cannot sign in before verifying their email.
19. A customer requesting another account's instance gets `404`, not `403`
no existence disclosure.
**Scoping:**
20. Every customer endpoint, called with a session for account A against a
resource of account B, returns `404`. Written as a table-driven test over the
route list so a new endpoint that forgets to scope fails the build.
## Verification before merge
1. Full suite green, including the scoping table test (test 20).
2. End to end, cloud: create account → create instance → issue Professional →
confirm the control-plane instance reports `valid` within 60 seconds with no
restart.
3. End to end, self-hosted: run `/setup` on a scratch install, copy the UUID,
link it, issue, download, paste, confirm `valid`.
4. Confirm admin's control-plane credential cannot write to `servers`, `keys` or
any collection other than `instances`.
5. Kill the admin service and confirm every Vantage instance keeps working
entirely normally.
## Risks
| Risk | Mitigation |
|---|---|
| Admin becomes a runtime dependency | Verification step 5; instances verify offline and never call admin |
| Signing key exposure | Single service, single variable, never in an image; rotation path from spec 1 |
| Cloud password now unlocks billing | Owner-only, rate-limited, audited, and stated in the release notes |
| Injection silently fails | Reconciliation every 15 minutes plus a staff health endpoint |
| Admin writes outside its remit in the control plane | Narrow code path; scoped Mongo credential; reviewed on every change |
| Self-hosted UUID squatted by another account | Unique index plus a non-disclosing error |
@@ -1,203 +0,0 @@
# Spec 4 — Admin Site
Date: 2026-07-24
Status: Design approved, not implemented
Depends on: spec 3 (`admin-backend`)
Ships: with spec 3. Can be developed in parallel with spec 5 once spec 3's API
is stable.
## Context
A fifth Next.js app, `adminsite/`, serving two audiences from one codebase:
- **Staff** — internal operators. Accounts, instances, licence history, plan
editing, injection health, audit.
- **Customers** — their own account, instances, licences and subscription state.
They share auth plumbing and a component library but almost no screens. The
split is by route group, so a customer route can never accidentally render a
staff view.
## Goals
1. A customer can buy, link a self-hosted instance, download a licence, and see
when it expires — without contacting anyone.
2. Staff can answer "why did this customer's instance stop working" in one screen.
3. Nothing about the marketing site or the control-plane UI changes.
## Non-goals
- Rebuilding billing management. Card details, invoices, payment methods and
cancellation all deep-link into Paddle's customer portal.
- Server management. This is not a second control plane; there is exactly one
link out to the instance and no data about servers, keys or workflows.
- Public signup for cloud. That stays on the marketing site (moving to admin's
backend in spec 5, but the *form* stays where customers already find it).
## Design
### App
Built exactly like `web/` and `site/`: Next.js 16 App Router, React 18, Tailwind
3, TanStack Query, `output: "standalone"`, `node:26-alpine`, listening on `3000`,
published as `3002`. In `docker-compose.site.yml` only.
`ADMIN_API_URL` is baked in at build time, as `API_URL` is for `web/`. It must be
**browser-reachable** and must appear in the backend's `ADMIN_ORIGIN`. Getting
this wrong is the single most common deployment failure in this repo's history —
`SITE_API_URL` has the same footgun documented in `CLAUDE.md` — so the app
renders an explicit "not connected" state rather than failing silently.
```
adminsite/
├── app/
│ ├── login/
│ ├── signup/ # self-hosted customer account creation
│ ├── verify/
│ ├── (customer)/
│ │ ├── page.tsx # account overview
│ │ ├── instances/[id]/
│ │ ├── instances/link/
│ │ ├── billing/
│ │ └── layout.tsx # customer nav, account guard
│ └── (staff)/staff/
│ ├── page.tsx # operations dashboard
│ ├── accounts/[id]/
│ ├── instances/[id]/
│ ├── licenses/
│ ├── plans/
│ └── layout.tsx # staff nav, staff guard
├── components/
└── lib/
```
Route-group layouts do the guarding. A customer session hitting `/staff/*` gets
redirected, not a 403 page — there is nothing to tell them about.
### Customer screens
**Overview** — the account, its instances as cards. Each card: name, cloud or
self-hosted, tier, licence state, expiry with days remaining, and a link either
to the instance's subdomain (cloud) or to its licence page (self-hosted).
Licence state is colour-coded and blunt: green valid, amber under 14 days, red
expired. An expired card says what still works — "servers and monitors are still
running; changes are disabled" — because that is the first thing a worried
customer wants to know.
**Instance detail** — tier, limits, features, subscription status, next renewal
date. For self-hosted: the linked UUID, a **Download licence** button, the blob
in a copy-to-clipboard box, and step-by-step paste instructions with the target
route named (`Settings → Licence` on their own install). A **Relink** action
showing the remaining allowance ("2 of 3 relinks remaining this term").
**Link an instance** — the self-hosted activation screen. Explains where to find
the UUID (shown on `/setup`, and permanently on `/settings/license`), takes the
paste, validates the format client-side, and on success issues the licence and
lands the customer directly on the download.
The whole flow — buy, link, download, paste — should be completable without
reading documentation. That is the bar for this screen.
**Billing** — subscription list with status and renewal date, plus a button to
Paddle's portal. Deliberately thin.
### Staff screens
**Dashboard** — the operational answers, not vanity metrics: licences expiring
in the next 14 days, subscriptions `past_due`, instances `awaiting_link` for more
than 48 hours, and **failed injections** from the reconciliation job. Each row
links straight to the thing that needs doing.
**Accounts** — searchable by name, email, Paddle customer ID and instance UUID.
Searching by UUID matters: a support email arrives containing a UUID and nothing
else.
**Account detail** — instances, subscriptions, customer users, audit trail.
**Instance detail** — everything about one instance, with the **full licence
history as a timeline**: issued, superseded, renewed, relinked, each with a
timestamp, reason and who did it. This is the screen that answers "why did this
stop working on the 14th". Actions: issue, reissue, relink without limit, and a
live view of the control-plane injection state for cloud instances.
**Licences** — global history, filterable by tier, deployment, expiry window and
issuance reason.
**Plans** — edit limits and features per tier. Two guard rails, because this
screen changes what every future customer gets:
- A confirmation step naming exactly what changes and stating that existing
licences are unaffected until reissued.
- The deployment field is not editable. Moving Free to `self_hosted` would break
the cloud-only rule that spec 1 leans on; changing it is a code review, not a
form field.
**Audit** — every mutating action, filterable.
### Design language
Visually distinct from `web/`. Staff regularly have both open, and a moment of
"which app am I in" before clicking Reissue is worth designing out. Different
accent colour and a persistent environment badge in the header (sandbox or
production, from a build-time flag) — clicking Issue against the wrong Paddle
environment should be hard.
Shared component patterns with `web/` where they exist; this is not a reason to
invent a second design system.
### Error and empty states
- Backend unreachable: a page-level "not connected" state naming
`ADMIN_API_URL`, matching the pattern the marketing site already uses.
- No instances yet: a customer-facing explanation of the two paths — buy cloud,
or buy self-hosted and link.
- `awaiting_link`: a prominent prompt on the overview, since a customer who has
paid and not linked is a customer who has paid for nothing yet.
- Licence download failure: show the blob inline as a fallback so the customer is
never blocked by a file download.
## Testing
Component and integration tests with mocked API responses. The repo has no
frontend test setup today; this is where one starts, scoped to the flows that
lose money or leak data when broken.
1. Customer session on `/staff/*` redirects; staff session reaches it.
2. Instance card renders correctly for each licence state, including expired,
and the expired copy names what still works.
3. Link flow: valid UUID succeeds and lands on download; malformed UUID is caught
client-side; already-linked UUID surfaces the backend's message.
4. Relink shows the remaining allowance and disables at zero with the support
message.
5. Licence download failure falls back to the inline blob.
6. Not-connected state renders when the API is unreachable.
7. Staff dashboard renders each alert category and links to the right resource.
8. Plan edit requires confirmation and shows the "existing licences unaffected"
wording.
9. Instance search by UUID returns the instance.
10. Licence history timeline renders every reason type in order.
## Verification before merge
1. Test suite green.
2. Full manual pass, self-hosted purchase to working licence, using only the UI
and no documentation — timed, and if it takes more than five minutes the flow
needs work.
3. Full manual pass, cloud: buy, confirm the licence appears in the control plane
within a minute, confirm the instance's own settings page agrees.
4. Staff pass: find an account by instance UUID, read its licence history,
reissue, confirm the control plane picks it up.
5. Responsive check at mobile width — a customer hit by an expiry email will open
this on a phone.
6. `docker build` from the repo root succeeds and the image runs with
`ADMIN_API_URL` baked in.
## Risks
| Risk | Mitigation |
|---|---|
| `ADMIN_API_URL` misconfigured at build | Explicit not-connected state; documented alongside the existing `SITE_API_URL` footgun |
| Staff action taken against the wrong environment | Persistent environment badge; confirmation on destructive actions |
| Customer confused by the self-hosted flow | Step-by-step link screen; five-minute bar in verification |
| Customer session reaching staff data | Route-group guards plus backend scoping (spec 3, test 20). Two layers, because one is not enough for this |
@@ -1,353 +0,0 @@
# Spec 2 — Instance Licensing and Enforcement
Date: 2026-07-24
Status: Design approved, not implemented
Depends on: spec 0a, spec 0b, spec 1 (`licensing-core`)
Ships: independently, with licenses issued by hand via `lkctl`. No admin site
needed.
## Context
Spec 1 defines what a license is. This spec makes the control plane hold one,
act on it, and let a self-hosted operator paste one in.
The guiding rule: **an expired license must never break a running fleet.** Agents
keep their keys, monitors keep watching, alerts keep firing. What stops is
growth and change. A customer whose card fails should be inconvenienced, not
paged at 3am because their monitoring went dark when Vantage decided to sulk.
## Goals
1. A license lives on the instance document and is verified on read.
2. Enforcement is deny-by-default: a new mutating route is gated because of where
it is mounted, not because someone remembered.
3. Degraded mode is obvious in the UI and reversible by pasting a valid license.
4. Self-hosted operators get an instance UUID they can hand to the admin site.
## Non-goals
- Issuing licenses. `lkctl` (spec 1) or the admin backend (spec 3).
- Any outbound network call. Verification is offline, permanently.
- Per-user or per-role licensing. The unit is the instance.
## Design
### Instance identity
Every install already has an `Instance` document with an `InstanceID` UUID. For
cloud instances this is created by signup; for self-hosted it is created by
`/setup`.
Change to `/setup`: after bootstrapping the first instance and its owner, the
setup page **displays the instance UUID** with a copy button and the text that
it is needed to activate a license. It is also shown permanently on
`/settings/license`.
No new identifier is invented. The instance UUID is the licensing identity.
### Storage
`shared/models.Instance` gains:
```go
LicenseBlob string `bson:"license_blob,omitempty" json:"-"`
LicenseTier string `bson:"license_tier,omitempty" json:"license_tier,omitempty"`
LicenseExpiry *time.Time `bson:"license_expiry,omitempty" json:"license_expiry,omitempty"`
```
The blob is authoritative. `LicenseTier` and `LicenseExpiry` are a denormalised
cache for listing and for the admin site's queries, rewritten from the verified
payload every time a blob is accepted. Nothing reads them for enforcement.
`LicenseBlob` is `json:"-"`. It is not a secret in the confidentiality sense —
it is signed public data — but there is no reason to spray it through API
responses.
### Runtime state
```go
type State struct {
Status license.State // valid | expired | invalid
Reason string
Tier string
ExpiresAt *time.Time
Limits license.Limits
Features map[string]bool
}
```
Resolved by `services.LicenseState(instanceID) State`, cached for 60 seconds
alongside the existing instance cache and invalidated immediately when a blob is
stored.
Three inputs, in precedence order:
1. `Instance.LicenseBlob`.
2. `VANTAGE_LICENSE` environment variable, used **only when the instance has no
stored blob**. This lets an automated self-hosted deployment ship a license
without a human pasting one. A blob stored through the UI always wins
afterwards, so an operator is never locked out by a stale environment value.
3. Neither → `Status: invalid`, `Reason: no_license`.
The verifier is called with `InstanceID` from the instance document and
`Deployment` from `VANTAGE_DEPLOYMENT` (`cloud` on our infrastructure,
`self_hosted` everywhere else, defaulting to `self_hosted`). The default matters:
an operator who removes the variable gets the stricter mode, not the looser one.
`invalid` and `expired` degrade identically. They differ only in the message.
### Enforcement
Three layers, deliberately separate because they answer different questions.
**Layer 1 — mutation gate.** A gin middleware `RequireActiveLicense` mounted on
the `/api` group, applying to every request whose method is not `GET` or `HEAD`.
```go
api := r.Group("/api", auth.RequireSession(), services.RequireActiveLicense())
```
Non-`valid``403 {"error":"license_required","state":"expired","reason":"..."}`.
Mounting at the group means **a route added tomorrow is gated by default**. That
is the whole point of putting it here rather than on individual handlers.
Explicit exemptions, allow-listed by path because they must work in degraded
mode:
| Route | Why |
|---|---|
| `POST /api/license` | Pasting a valid license is how you recover |
| `POST /auth/*` | Login and logout are outside `/api` already; listed for clarity |
| `DELETE` on any resource | Deleting is how you get back under a limit |
| `POST /api/servers/:id/apply-updates` | Security patching must never be paywalled |
The `DELETE` exemption deserves emphasis: a customer downgraded to Free with 10
servers must be able to remove 7 of them. Blocking deletes would trap them.
**Layer 2 — feature gate.** `RequireFeature(name)` on the route groups that need
it:
- `console``POST /api/console/connect`, `GET /api/console/tunnel`
- `oidc``GET,PUT /api/instance/oidc`
Missing feature → `403 {"error":"feature_unavailable","feature":"console"}`.
OIDC needs care: `/auth/oidc/start` and `/auth/oidc/callback` are unauthenticated
and outside `/api`. They check the feature directly and, if unavailable, redirect
to `/login?error=oidc_unavailable` rather than returning JSON. **Existing OIDC
sessions are not terminated** — losing the feature stops new SSO logins, it does
not evict people mid-session.
**Layer 3 — limits.** Enforced in the service layer, because a limit needs a
count that middleware does not have:
| Limit | Checked in |
|---|---|
| `max_servers` | `services.CreateServer` / `POST /api/servers/new` |
| `max_secret_groups` | `services.CreateSecretGroup` |
| `max_channels` | `services.CreateChannel` |
`-1` means unlimited. Over limit → `403 {"error":"limit_exceeded","limit":"max_servers","current":3,"max":3}`.
Counts are of live rows: revoked assignments and deleted servers do not count.
**Over-limit instances are never truncated.** A Professional instance with 20
servers that lapses to Free keeps all 20 running; it simply cannot add a 21st.
Deleting resources is always permitted. Silently disabling a customer's servers
because their card expired is not a behaviour this system will have.
### Background work in degraded mode
This is where "read-only" needs to be specific, because these paths do not go
through gin at all.
| Subsystem | Degraded behaviour |
|---|---|
| **Monitor scheduler** | **Keeps running.** Checks execute, incidents open, notifications fire. |
| Monitor create/edit/delete | Blocked by layer 1 (delete exempted). |
| Workflow runner | New runs blocked by layer 1. **In-flight runs finish** rather than being killed mid-step — a half-run workflow is worse than a completed one. |
| Agent `SyncKeys` | Returns the existing desired key set unchanged. Nothing is torn off disk. New assignments cannot be created, so nothing changes anyway. |
| Agent registration | A **new** agent registering against an over-limit instance is refused with a clear message; existing agents re-register freely. |
| Inventory, heartbeat, update reporting | Unaffected. |
| `ApplyUpdatesCmd` | Allowed. Security patching is not gated. |
| ESO secrets read (`GET /api/secrets/:group/values`) | **Allowed.** It is a `GET`, and breaking a Kubernetes cluster's secret sync over a billing state is disproportionate. |
| Log retention sweep, offline sweep | Unaffected. |
Keeping monitors alive is a deliberate reversal of a stricter earlier draft. It
is the single most important line in this spec: **billing state must not take
away a customer's ability to know their infrastructure is on fire.**
### API
```
GET /api/license any authenticated user
POST /api/license owner only
```
`GET` returns:
```json
{
"instance_id": "…",
"state": "valid",
"reason": "",
"tier": "professional",
"expires_at": "2027-07-24T00:00:00Z",
"days_remaining": 365,
"limits": { "max_servers": -1, "max_secret_groups": -1, "max_channels": -1 },
"features": { "console": true, "oidc": true },
"usage": { "servers": 12, "secret_groups": 4, "channels": 2 },
"source": "stored"
}
```
`usage` is included so the UI can render "12 of 3 servers" honestly when an
instance is over its limit, rather than pretending.
`POST` takes `{"blob": "..."}`, verifies with the instance's own ID and
deployment mode, and on success stores the blob, refreshes the cache, and writes
an audit event. On failure it returns `400` with the specific reason:
| Reason | Message |
|---|---|
| `bad_signature` | This licence key is not valid. Check it was copied in full. |
| `deployment_mismatch` | This licence is for Vantage Cloud and cannot be used on a self-hosted install. |
| `instance_mismatch` | This licence was issued for a different instance. Your instance ID is `<uuid>`. |
| `expired` | This licence expired on `<date>`. |
An **expired** blob is still stored if it is otherwise valid, so the UI can show
what expired and when. An **invalid** blob is rejected and the previous one kept.
Rate-limited to 10 attempts per instance per hour. There is no oracle here worth
protecting, but an unbounded verify endpoint is an unbounded CPU endpoint.
### Frontend
`useLicense()` hook over `GET /api/license`, cached by TanStack Query and
invalidated after a successful paste.
- **Banner, persistent, top of every page** when `state != valid`:
- `expired` — "Your Vantage licence expired on `<date>`. Your servers and
monitors are still running, but changes are disabled until it is renewed."
with a link to the admin site.
- `invalid` / `no_license` — "This instance has no valid licence. Add one in
Settings → Licence."
- **Warning banner** in the final 14 days of a valid term, dismissible per
session.
- **Gated features render disabled with an upgrade tooltip, not hidden.** A
customer cannot buy what they cannot see, and a feature that vanishes reads as
a bug.
- **Limit indicators** on the servers, secrets and channels list pages: "3 of 3
servers used" with the create button disabled at the cap.
- `/settings/license`: current state, tier, expiry, limits with live usage, the
instance UUID with a copy button, and a textarea plus file upload for a new
blob. Owner-only; other roles see the state read-only.
### Grandfathering existing tenants
Migration `0005_grandfather_licenses`, cloud only, guarded on
`VANTAGE_DEPLOYMENT == "cloud"`:
For every instance with no `license_blob`, issue a Professional license expiring
**one year** from the migration date and store it.
The migration cannot sign — the server has no private key and, per spec 1, no
signing code. So the blobs are **generated ahead of time with `lkctl`** and
supplied to the migration through `VANTAGE_GRANDFATHER_BLOBS`, a JSON map of
instance ID to blob. The migration stores what it is given, verifies each blob
against its instance before storing, and logs any instance it had no blob for.
Clumsy, and correct. The alternative is putting a signing key in the control
plane, which is the thing this design most wants to avoid.
Self-hosted installs are not grandfathered. On upgrade they land in `no_license`
and read-only until an operator pastes a key — which is the intended behaviour
for a paid product, and is why the release notes must lead with it.
## Testing
**Unit, no database:**
1. `State` resolution precedence: stored blob wins over `VANTAGE_LICENSE`;
environment used when no blob; neither → `no_license`.
2. Feature map construction from the payload's `Features` slice.
3. Limit comparison with `-1`, with zero, and with a count exactly at the cap.
**Middleware, with a stub state:**
4. `GET` passes in every state.
5. `POST`/`PUT`/`DELETE` pass when `valid`, fail `403` when `expired` and when
`invalid` — except `DELETE`, which passes in all states.
6. `POST /api/license` passes when `expired` (the recovery path).
7. `POST /api/servers/:id/apply-updates` passes when `expired`.
8. `RequireFeature("console")` passes with the feature, `403`s without it.
9. **Coverage test:** enumerate every registered route and assert that every
non-`GET` route is either behind `RequireActiveLicense` or on the exemption
allow-list. This test is what stops layer 1 rotting as routes are added.
**Service layer, against MongoDB:**
10. `CreateServer` at the cap → `limit_exceeded`; one below → succeeds.
11. Over-limit instance can still `DELETE` a server, and can create again once
back under the cap.
12. Deleted and revoked rows do not count toward limits.
**Degraded background behaviour:**
13. Monitor scheduler executes checks for an instance with an expired license.
14. An incident opened during degraded mode still dispatches notifications.
15. `SyncKeys` for an expired instance returns the same key set as before expiry.
16. A new agent registering against an over-limit instance is refused; an
existing agent re-registers successfully.
17. A workflow run in flight when the license expires completes its remaining
steps.
**API:**
18. `POST /api/license` with a valid blob stores it and flips state to `valid`.
19. Each rejection reason returns its own message and leaves the stored blob
untouched.
20. An expired-but-well-formed blob is stored and reported as `expired`.
21. Non-owner `POST``403`.
**Migration:**
22. `0005` stores and verifies supplied blobs, skips instances that already have
one, logs instances with no blob supplied, and is a no-op when
`VANTAGE_DEPLOYMENT != "cloud"`.
## Verification before merge
1. Full test suite green, including the route-coverage test (test 9).
2. Manual pass on a scratch instance: issue a Professional license with `lkctl`,
paste it, confirm everything works. Issue one expiring in 60 seconds, wait,
confirm the banner appears, mutations `403`, **monitors keep firing**, and
pasting a fresh license restores normal operation without a restart.
3. Manual pass on the Free tier: confirm the 3-server cap, that console and OIDC
are visibly disabled with upgrade tooltips, and that a 4th server is refused
with a clear message.
4. Confirm a cloud-issued Free license is rejected on a `self_hosted` install
with `deployment_mismatch`.
5. Confirm a license issued for another instance is rejected with
`instance_mismatch` and the message shows the correct local UUID.
## Rollout
1. Generate grandfather blobs with `lkctl` for every existing cloud instance.
2. Deploy with `VANTAGE_GRANDFATHER_BLOBS` set; migration `0005` runs.
3. Verify every cloud instance reports `valid`, Professional, one year out.
4. Unset the variable on the next deploy — it is single-use.
5. Release notes for self-hosted must state plainly that upgrading requires a
licence key, and how to get one.
## Risks
| Risk | Mitigation |
|---|---|
| A mutating route added later without a gate | Route-coverage test (test 9) fails the build |
| Customer locked out and unable to recover | `POST /api/license` and all `DELETE`s exempt from the gate |
| Existing cloud tenants degrade on deploy | Migration 0005, verified before the traffic switch |
| Over-limit customer trapped | Deletes always allowed; existing resources never truncated |
| Clock wrong on a self-hosted host | `Verify` warns on a future `IssuedAt`; documented in the licence settings page |
| Monitoring lost on billing failure | Explicitly designed out — the scheduler ignores licence state |
@@ -1,240 +0,0 @@
# Spec 0b — Org to Instance Rename
Date: 2026-07-24
Status: Design approved, not implemented
Depends on: spec 0a (`shared-module`)
Ships: independently, before any licensing code
## Context
The licensing model separates two concepts that the codebase currently conflates
under one word:
- **Account** — a paying customer. Lives only in the admin control plane
(spec 3). The control plane never learns about it.
- **Instance** — one deployment of Vantage: its own subdomain, its own users,
its own servers, keys, workflows, monitors and secrets. One license attaches
to one instance.
Today's control-plane `Org` **is** an Instance. An Account may hold several,
some cloud and some self-hosted, and the self-hosted ones have no row in the
cloud database at all.
Keeping the name `Org` would leave the control plane using a word that means
something different in the admin site, in Paddle, and in every support
conversation. This spec renames it everywhere, including on disk.
This is the highest-risk change in the programme: `org_id` is the tenant
isolation key on every document in every collection. It is done alone, before
anything else, so that nothing else is in flight when it deploys.
## Goals
1. `Instance` is the only word for a tenant, in code, API, UI and database.
2. No document is lost and no tenant scoping is weakened.
3. The migration is reversible.
## Non-goals
- Any behaviour change. Same routes' semantics, same permissions, same data.
- Introducing Accounts. The control plane never gets them.
- Touching the agent. It talks gRPC and has no concept of a tenant.
## Design
### Naming map
| Today | After |
|---|---|
| collection `orgs` | `instances` |
| collection `org_oidc` | `instance_oidc` |
| field `org_id` (all collections) | `instance_id` |
| `models.Org` | `models.Instance` |
| `Org.OrgID` | `Instance.InstanceID` |
| `User.OrgID`, `Settings.OrgID`, every `OrgID` field | `InstanceID` |
| `services/orgs.go`, `GetOrg`, `CreateOrg`, `ListOrgIDs`, `CountOrgs`, `FirstOrg`, `AdoptOrg`, `GetOrgBySlug` | `services/instances.go`, `GetInstance`, `CreateInstance`, … |
| `services/org_oidc.go` | `services/instance_oidc.go` |
| `auth/orghost.go` | `auth/instancehost.go` |
| `/api/org/users`, `/api/org/oidc` | `/api/instance/users`, `/api/instance/oidc` |
| `shared/provision.CreateOrg`, `RollbackOrg` | `CreateInstance`, `RollbackInstance` |
| session field `org_id` | `instance_id` |
| `GET /auth/me` response `org_id` / `org` | `instance_id` / `instance` |
| UI copy "Organisation" | "Instance" |
Reserved slugs gain no new entries here, but note `admin` is already reserved,
which the admin site relies on later.
### Collections carrying `org_id`
All of: `servers`, `keys`, `assignments`, `users`, `org_oidc`, `settings`,
`secrets`, `workflows`, `workflow_steps`, `workflow_runs`, `monitors`,
`incidents`, `monitor_rollups`, `notification_channels`, `console_sessions`,
`audit_logs`, plus `orgs` itself. `migrations` does not carry one.
`site_pending_signups` does not carry `org_id`, but its `org_name` field becomes
`instance_name` for consistency; it is sitesvc-private so this is free.
The migration must derive this list from a constant in code, not from a
hand-written list in a runbook, so that a collection added between design and
deploy is not silently missed:
```go
var scopedCollections = []string{ /* the list above */ }
```
A boot-time assertion (spec 2 onwards) checks that no collection outside this
list contains an `org_id` field. Cheap insurance against a future collection
being added without being renamed.
### Migration `0004_org_to_instance`
Recorded in `migrations` like the existing three. Runs after
`0003_missed_org_scopes`.
**The migration only renames. It never deletes and never drops.** A bad deploy
is recovered by running the inverse rename, not by restoring a backup.
Steps, in order:
1. **Guard.** If collection `instances` already exists and `orgs` does not, the
migration has already run against this database by an earlier binary; record
the marker and return. Idempotency matters because the marker write and the
data work are not in one transaction.
2. **Rename collections.** `orgs``instances`, `org_oidc``instance_oidc`,
via `adminCommand{renameCollection}`. Fails loudly if the target exists.
3. **Rename the field.** For each collection in `scopedCollections`:
`UpdateMany({org_id: {$exists: true}}, {$rename: {"org_id": "instance_id"}})`.
Record `matched` and `modified` per collection in the log.
4. **Verify.** For each collection, assert
`CountDocuments({org_id: {$exists: true}}) == 0` and
`CountDocuments({instance_id: {$exists: true}}) == totalCount`. Any mismatch
aborts before the marker is written, leaving the migration to retry.
5. **Indexes.** Drop and recreate indexes that name `org_id` in their key spec:
unique `settings.instance_id`, the ESO token-hash index, and any compound
scoping indexes. Unique `instances.slug` and `users.email` are unaffected by
the field rename but are re-declared idempotently.
6. **Write the marker.**
Steps 24 are not atomic across collections. Mongo multi-document transactions
would require a replica set, which is not guaranteed for self-hosted installs.
Instead the migration is written to be **safely re-runnable**: `$rename` on a
document that has already been renamed matches nothing, and the collection
rename is guarded in step 1.
Rollback, if ever needed, is the same code with the rename reversed, shipped as
a one-shot command rather than a migration — deliberately manual, because the
only reason to run it is a decision to revert the release.
### Version skew
`sitesvc` and `server` write the same documents. A skew where one writes
`org_id` and the other reads `instance_id` creates tenants that are invisible to
the application — the exact failure `CLAUDE.md` warns about.
After spec 0a both read the shape from `shared`, so the skew window is a
deployment-ordering problem rather than a code-drift problem:
- Both images are built from the same commit and deployed together.
- The migration runs from the `server` container at boot, as the existing three
do.
- `sitesvc` at boot asserts that collection `instances` exists and refuses to
start otherwise, with the message
`instances collection not found; deploy the control plane first`. Failing to
start is strictly better than provisioning into a collection nobody reads.
The self-hosted deployment runs no sitesvc, so it sees only the server change.
### API and frontend
REST route renames are **breaking**, but every consumer is first-party (`web/`)
and ships in the same release. No compatibility aliases — a permanent dual path
in the tenant-scoping layer is worse than a coordinated release.
`web/` changes: the API client's paths, the `useMe` shape, all UI copy from
"Organisation" to "Instance", and the settings route `/settings/org`
`/settings/instance`.
`site/` marketing copy changes where it says "organisation" about a tenant. Where
it means the customer, it becomes "account" — that word now has a specific
meaning and the marketing site is the first place a customer meets it.
## Testing
**No automated tests.** Decision taken 2026-07-24, consistent with spec 0a.
This is the change where that costs the most: it moves the tenant isolation key
across 17 collections, and a mistake orphans a customer's entire fleet rather
than breaking a build. The compensating controls are therefore not optional, and
the implementation plan makes each a mandatory step:
1. **Dry run against a restored copy** before the code is even committed —
migrate a `mongorestore`d duplicate of production and read the per-collection
rename counts.
2. **Idempotency by hand** — run the dry run twice; the second must complete
with no error and nothing left to rename.
3. **Interrupted-run recovery by hand** — rename `orgs` manually, then run the
migration; it must complete and leave every document carrying `instance_id`.
4. **Count comparison against a production snapshot** — record every
collection's document count before and after; any difference stops the
release.
5. **Per-tenant isolation comparison** — for three real tenants, count rows in
`servers`, `keys`, `workflows`, `monitors`, `secrets` and `audit_logs` by
`org_id` before and by `instance_id` after. Identical, or the release stops.
This is the check that proves tenant isolation survived.
6. **Stale-field sweep** — assert no collection anywhere still holds an
`org_id`.
7. **Rollback rehearsal** — migrate a third copy, run `rename-rollback`, confirm
the counts return to baseline and the pre-release binary boots against it.
Deploying without having done this is not permitted.
8. **Boot guard, both directions** — sitesvc must refuse an unmigrated database
and start normally against a migrated one.
`AssertNoScopedCollectionMissed` runs at every boot and is fatal. With no test
suite it is the standing protection against a future collection being added
without being added to `ScopedCollections`.
## Verification before merge
Run against a **restored production snapshot**, not a synthetic database:
1. Record `db.getCollectionNames()` and per-collection `countDocuments()` before.
2. Run the migration.
3. Assert every count is identical afterwards.
4. Assert `instances.countDocuments()` equals the old `orgs.countDocuments()`.
5. Pick three real tenants; run the same scoped query before (by `org_id`) and
after (by `instance_id`) and confirm identical result sets. This is the test
that proves tenant isolation survived.
6. Boot the server against the migrated snapshot; log in as a real user; confirm
servers, keys, workflows, monitors and secrets all list correctly.
7. Boot sitesvc against the migrated snapshot; complete a signup end to end.
8. Boot sitesvc against an **un**migrated snapshot; confirm it refuses to start
with the expected message.
## Rollout
1. Take a database backup. Not optional — this is the one change where the
inverse rename is the recovery path and the backup is the second.
2. Deploy `server`, `web`, `site` and `sitesvc` from one commit, together.
3. Server boots, migration runs, marker recorded.
4. Watch for the sitesvc guard message; if it appears, sitesvc started first and
will restart cleanly.
Expect a short window during the server restart where the API is unavailable.
Agents are unaffected: they reconnect, and no gRPC message carries a tenant ID.
## Risks
| Risk | Mitigation |
|---|---|
| Partial migration leaves mixed field names | Step 4 verification aborts before the marker; migration is re-runnable |
| A collection missed from the list | List is a code constant plus a completeness test plus a boot-time assertion |
| sitesvc deployed before server | Boot guard refuses to start |
| An index still keyed on `org_id` | Step 5 drops and recreates; verification includes an index listing diff |
| A hard-coded `org_id` string outside the model layer | `grep -rn '"org_id"' server/ sitesvc/ shared/` must return only the migration file after the change |
| Frontend missed a renamed route | Full manual pass over every route in the UI before release |
## Follow-on
With `Instance` established, spec 1 (`licensing-core`) can define a license
payload that binds to `instance_id` without inventing a word the codebase does
not use.
@@ -1,294 +0,0 @@
# Spec 1 — Licensing Core
Date: 2026-07-24
Status: Design approved, not implemented
Depends on: spec 0a (`shared-module`), spec 0b (`instance-rename`)
Ships: independently. Adds a package and a CLI; changes no running behaviour.
## Context
Licenses are **offline-verified signed blobs**. A Vantage server checks a
signature and an expiry date and asks nobody's permission. That choice buys
self-hosted installs that work in air-gapped networks and a control plane with no
licensing availability dependency.
It costs revocation. Once issued, a license is valid until it expires, whatever
Paddle later says. Every other decision in the programme follows from accepting
that: Self Hosted is annual-only so the unenforceable window is bounded, and
cancellation takes effect at term end rather than immediately (spec 5).
This spec defines the payload, the signing and verification, and a CLI to issue
licenses by hand. It deliberately lands before the admin site so that specs 1+2
together give working licensing with no new service to operate.
## Goals
1. One struct, in `shared`, read identically by the verifier and the issuer.
2. Verification that needs no network, no clock sync beyond a rough one, and no
configuration.
3. A hand-issuance path good enough to run production on until spec 3 lands.
## Non-goals
- Storing licenses. Spec 2 owns the instance document; spec 3 owns issuance
history.
- Deciding tier contents. Tiers are data; the values in this spec are the
initial seed, and spec 3's `plans` table becomes their home.
- Any phone-home, revocation list or online check. There is none, anywhere, by
design.
## Design
### Package
`shared/license/`, inside the module created by spec 0a:
```
shared/license/
├── license.go # License, Limits, feature constants
├── sign.go # Sign, build-tagged out of the server binary
├── verify.go # Verify, Parse
├── keys.go # trustedPublicKeys
└── license_test.go
```
Uses `github.com/hyperboloide/lk` (ECDSA P-384 with SHA-256, base32 encoding).
### Payload
```go
package license
type License struct {
ID string `json:"id"` // uuid, for support and audit
InstanceID string `json:"instance_id"` // the instance this license is bound to
AccountID string `json:"account_id"` // admin-side customer, informational
InstanceName string `json:"instance_name"` // display only
Tier string `json:"tier"` // "free" | "professional" | "self_hosted"
Deployment string `json:"deployment"` // "cloud" | "self_hosted"
IssuedAt time.Time `json:"issued_at"`
ExpiresAt time.Time `json:"expires_at"`
Limits Limits `json:"limits"`
Features []string `json:"features"`
}
type Limits struct {
MaxServers int `json:"max_servers"` // -1 means unlimited
MaxSecretGroups int `json:"max_secret_groups"`
MaxChannels int `json:"max_channels"`
}
const (
FeatureConsole = "console" // browser SSH/RDP/VNC
FeatureOIDC = "oidc" // per-instance single sign-on
)
const (
TierFree = "free"
TierProfessional = "professional"
TierSelfHosted = "self_hosted"
DeploymentCloud = "cloud"
DeploymentSelfHosted = "self_hosted"
)
```
`InstanceID` is **always populated**. There is no unbound license: the
self-hosted purchase flow (spec 4) links the instance UUID before the license is
issued, so binding happens at signing time. This removes the claim endpoint, the
best-effort phone-home and the multi-claim reconciliation that an unbound design
would have needed.
**The server never branches on `Tier`.** It reads `Limits` and `Features` only.
`Tier` exists for display, support and analytics. Adding a tier, or changing what
a tier includes, must never require a server release.
### Tier seed values
Recorded here as the initial contents of spec 3's `plans` table. Snapshotted into
each license at issue, so changing the table never rewrites an issued license —
the same principle as `workflow_runs.steps_snapshot`.
| | Free | Professional | Self Hosted |
|---|---|---|---|
| `deployment` | `cloud` | `cloud` | `self_hosted` |
| `max_servers` | 3 | -1 | -1 |
| `max_secret_groups` | 1 | -1 | -1 |
| `max_channels` | 1 | -1 | -1 |
| `console` | no | yes | yes |
| `oidc` | no | yes | yes |
| billing term | monthly, £0 | monthly or annual | **annual only** |
Free is cloud-only. A self-hosted install can never hold a valid Free license
because Free is only ever signed with `deployment: "cloud"`, and verification
rejects a deployment mismatch. There is no server-side flag to edit.
### Signing
```go
//go:build !noSign
func Sign(l License, privateKeyHex string) (string, error)
```
Marshals to canonical JSON, signs with lk, returns the base32 blob.
`Sign` is excluded from the server binary with a build tag. The server has no
reason to hold signing code and there is no reason to ship it into a customer's
data centre.
The private key lives in `LICENSE_SIGNING_KEY` on the issuing side only — the
CLI now, the admin backend from spec 3. It is never in the repo, never in an
image, never in the control plane's environment.
### Verification
```go
type VerifyOpts struct {
InstanceID string // required: the verifier's own instance
Deployment string // required: "cloud" or "self_hosted"
Now time.Time // injectable for tests
}
type Result struct {
License License
State State // Valid, Expired, Invalid
Reason string
}
const (
StateValid State = "valid"
StateExpired State = "expired"
StateInvalid State = "invalid"
)
func Verify(blob string, opts VerifyOpts) Result
```
Checks, in order, stopping at the first failure:
1. Blob decodes and the signature verifies against one of `trustedPublicKeys`.
Failure → `Invalid`, reason `bad_signature`.
2. `l.Deployment == opts.Deployment`. Failure → `Invalid`, reason
`deployment_mismatch`. This is the check that makes Free cloud-only.
3. `l.InstanceID == opts.InstanceID`. Failure → `Invalid`, reason
`instance_mismatch`.
4. `opts.Now.Before(l.ExpiresAt)`. Failure → `Expired`.
5. Otherwise `Valid`.
**`Expired` and `Invalid` are distinct states and the caller treats them
differently in messaging** (spec 2), even though both degrade the instance the
same way. A customer whose card failed and a customer who pasted the wrong blob
need different words.
`Parse(blob) (License, error)` verifies the signature only, ignoring binding and
expiry. Used by the admin site to display a license and by support to inspect a
blob a customer has emailed in. Never used for enforcement.
Clock skew: no tolerance is applied. Terms are a month or a year; a server whose
clock is wrong by enough to matter has bigger problems, and a tolerance window is
a thing to get wrong. `Verify` logs at warn level if `IssuedAt` is in the future,
which is the signal that a clock is badly off.
### Key management
```go
// trustedPublicKeys is ordered. Index 0 is the current signing key.
// To rotate: prepend the new key, ship a server release, then reissue.
// Remove a retired key only after every license signed with it has expired.
var trustedPublicKeys = []string{
"<base32 ECDSA P-384 public key>",
}
```
A slice from day one even though it holds one entry, because retrofitting a
single-key verifier into a multi-key one during an incident is not a thing to
plan for.
Public keys are compiled in. They are not configurable, because a configurable
trust root is a licensing bypass: a self-hosted operator could point it at a
keypair they generated.
Key generation is a documented one-off:
```
go run ./shared/license/cmd/lkgen keypair
```
prints a private key (base32) for the vault and a public key (base32) to paste into
`keys.go`. The private key is stored in a password manager and in the admin
service's environment. **If it is lost, no new licenses can be issued for any
existing customer without a server release.** Back it up in two places.
### CLI issuer
`shared/license/cmd/lkctl`, built only for internal use:
```
lkctl keypair
lkctl issue --instance-id=<uuid> --instance-name="Acme" \
--tier=professional --deployment=cloud \
--term=1y [--account-id=<id>] [--out=acme.lic]
lkctl inspect <file-or-blob>
```
`issue` reads `LICENSE_SIGNING_KEY`, applies the tier seed values from a table
compiled into the CLI, and prints the blob. `--term` accepts `1m`, `1y` or an
explicit `--expires=RFC3339`.
This is the production issuance path until spec 3 ships. It is kept afterwards
for support and disaster recovery — if the admin service is down and a customer's
license expires, a blob can still be cut by hand.
Issued blobs from `lkctl` are not recorded anywhere. Spec 3 backfills its
`licenses` table from `inspect` output when it takes over.
## Testing
`shared/license` is pure and needs no database, so this suite is fast and
thorough. Written test-first.
1. Round trip: `Sign` then `Verify` returns `Valid` with an identical payload.
2. Tampering: flip one character of the blob → `Invalid`, `bad_signature`.
3. Tampering with intent: re-sign a payload with a *different* keypair →
`Invalid`. This is the test that proves an attacker cannot mint licenses.
4. Expiry: `ExpiresAt` one second in the past → `Expired`. One second in the
future → `Valid`.
5. Deployment mismatch: a Free (`cloud`) license verified with
`Deployment: "self_hosted"``Invalid`, `deployment_mismatch`.
6. Instance mismatch: correct signature, different `InstanceID``Invalid`,
`instance_mismatch`.
7. Check order: a blob that is both expired *and* instance-mismatched reports
`instance_mismatch`, not `Expired`. Order is part of the contract because the
reason drives the message.
8. Multi-key: a license signed with `trustedPublicKeys[1]` verifies. One signed
with a key not in the slice does not.
9. `Parse` returns the payload for an expired and for a mismatched license, and
errors for a bad signature.
10. Unicode and long instance names survive the round trip.
11. Golden blob: a fixture blob checked into the repo, signed with a **test-only**
keypair, must keep verifying. This catches an accidental change to the
canonical JSON encoding, which would silently invalidate every issued
license in the field.
Test 11 matters more than it looks. The encoding is part of the wire format.
## Verification before merge
1. `go test ./shared/license/...` passes, including the golden fixture.
2. `lkctl keypair``lkctl issue``lkctl inspect` round trips at the command
line.
3. `go build -tags noSign ./server/...` succeeds and
`go tool nm` on the resulting binary shows no `license.Sign` symbol.
4. The production keypair is generated, the private half stored in two places,
and the public half committed in `keys.go`.
## Risks
| Risk | Mitigation |
|---|---|
| Signing key lost | Documented two-location backup; generation is a one-off with an explicit checklist |
| Signing key leaked | Rotation path exists from day one: prepend key, release, reissue. Retire the old key once its licenses expire |
| Canonical encoding changes | Golden fixture test |
| Signing code shipped to customers | Build tag plus a symbol check in verification |
| No revocation | Accepted and documented. Bounded by term length; Self Hosted is annual-only |
@@ -1,271 +0,0 @@
# Spec 5 — Paddle Billing
Date: 2026-07-24
Status: Design approved, not implemented
Depends on: spec 3 (`admin-backend`)
Ships: after spec 3. Can be developed in parallel with spec 4.
## Context
Paddle is merchant of record: it owns checkout, tax, invoices, dunning and the
customer billing portal. This spec connects Paddle's subscription lifecycle to
the licence issuance functions spec 3 defines, and moves cloud signup off
sitesvc.
The central constraint, restated because every table below follows from it:
**licences are offline-verified, so nothing Paddle says can revoke one early.**
Cancellation takes effect when the licence expires. Self Hosted is annual-only to
bound that window; the alternative — a customer holding a valid key for eleven
months after cancelling a monthly plan — is not acceptable.
## Goals
1. A catalog in Paddle sandbox, promotable to production by configuration alone.
2. Webhooks that issue and renew licences reliably, including under retries and
out-of-order delivery.
3. Cloud signup owned by one service instead of two.
## Non-goals
- Building any part of billing Paddle already provides.
- Usage-based or metered pricing. Tiers are flat.
- Proration logic. Paddle handles money; we react to the resulting subscription
state.
## Design
### Catalog
Three products, created in **sandbox** first. Production is a configuration
change: the same `plans` rows carry different `paddle_product_id` and
`paddle_price_ids`, selected by `PADDLE_ENV`.
| Product | Prices | Notes |
|---|---|---|
| Vantage Free | monthly, £0 | Yes, a real £0 subscription. It gives every account a Paddle customer, a lifecycle, and an upgrade path with no special-case code. |
| Vantage Professional | monthly, annual | Cloud |
| Vantage Self Hosted | **annual only** | No monthly price exists, so the offline-revocation window is at most a year |
**No price ID is ever hard-coded.** They live in `plans.paddle_price_ids` and are
edited through the staff UI. A price change in Paddle is a data edit, not a
deploy.
`custom_data` on every checkout carries `{ account_id, instance_id, tier }`. This
is what lets a webhook route without a lookup table, and it is why the
self-hosted flow creates the instance record *before* checkout completes.
### Checkout
Paddle Checkout, overlay mode, in the admin site.
**Cloud upgrade** — instance exists, `instance_id` in `custom_data`, existing
Paddle customer reused.
**Self-hosted purchase** — the instance does not exist yet. Order:
```
Customer creates an admin-site account (verified email)
Account row created, then an admin_instances row with status awaiting_link
and a generated placeholder instance record
Checkout opened with account_id and that instance row's id in custom_data
subscription.created fires → subscription recorded, status awaiting_link,
NO licence issued
Customer pastes their install's UUID → instance_id set, status active
→ licence issued and delivered
```
The instance row exists before payment so the webhook has something to attach to.
The licence is not issued until the UUID is known, because a licence with no
instance to bind to cannot be signed — spec 1 has no unbound licence.
A customer who pays and never links has a subscription and no licence. Spec 4's
staff dashboard flags `awaiting_link` older than 48 hours, and a reminder email
goes out at 24 hours and 72 hours. This is the most likely place for a paying
customer to get stuck, so it gets active chasing rather than a support queue.
### Webhooks
`POST /api/paddle/webhook`, signature-verified with `PADDLE_WEBHOOK_SECRET`.
An unsigned or badly signed request is rejected `401` and logged — never
processed.
**Idempotency is mandatory.** Paddle retries. Every event ID is recorded in
`paddle_events` with a unique index before processing; a duplicate returns `200`
without acting. `200` on duplicates matters — returning an error would make
Paddle retry a message we have already handled, forever.
| Event | Action |
|---|---|
| `subscription.created` | Record the subscription. Cloud: issue and inject. Self-hosted: leave `awaiting_link`, issue nothing. |
| `subscription.updated` | Tier or term changed: issue a replacement licence at the new tier, supersede the old. Cloud injects; self-hosted emails a new blob and flags the site. Reflects Paddle's resulting state; no proration maths here. |
| `subscription.canceled` | Mark `cancelled`. **No licence action.** The current licence runs to expiry, then the instance degrades per spec 2. |
| `subscription.past_due` | Mark `past_due`, notify the customer, flag for staff. Licence untouched. Dunning is Paddle's job; ours is not to punish a retryable card failure. |
| `transaction.completed` where the transaction is a subscription renewal | Issue the next term's licence, supersede, inject or email. Reset `RelinkCount`. |
| `transaction.payment_failed` | Record for staff visibility. No licence action. |
| `customer.updated` | Sync `billing_email` onto the account. |
Out-of-order delivery is handled by making every handler a function of the
subscription's *current* state as reported in the event payload, rather than of
the transition. An `updated` arriving before its `created` creates the
subscription row and proceeds.
Renewal licences are issued with a **3-day grace** past the period end (spec 3),
so a webhook delayed by hours never produces a gap in coverage.
**Webhook failures must be visible.** Every failed handler writes to
`admin_audit` and appears on the staff dashboard. A licence that silently failed
to issue is a customer who paid and got nothing.
### Cancellation, stated plainly
When a customer cancels:
- Paddle stops billing at period end.
- We issue no further licences.
- Their current licence keeps working until it expires — up to a month for
Professional monthly, up to a year for Self Hosted.
- On expiry the instance degrades per spec 2: monitors keep running, changes stop.
This is documented in the terms and shown on the cancellation confirmation
screen, because a customer who cancels and sees their instance keep working
should understand why rather than assume the cancellation failed.
### Signup migration off sitesvc
Cloud signup currently lives in sitesvc: `site_pending_signups`, a verification
email, and provisioning on link click. It now needs to also create an Account, a
Paddle customer, a Free subscription and a licence.
**Signup moves to the admin backend.** The form stays on the marketing site where
customers find it, but it posts to admin instead of sitesvc. sitesvc keeps the
contact form only.
The reason is the one `CLAUDE.md` already names: provisioning logic duplicated
across services drifts. Spec 0a removed the second copy; adding signup to admin
while leaving it in sitesvc would create a third.
New flow, preserving every property of the current one:
```
Marketing site form → POST /api/signup on admin
→ pending record, password bcrypt cost 12, token 32 random bytes,
only the SHA-256 hash stored, 24h expiry, TTL index
→ verification email
Link opened → FindOneAndDelete the pending record (atomic, before provisioning)
→ shared.CreateInstance + shared.CreateUser in the control plane
→ Account created
→ Paddle customer created, Free subscription created
→ Free licence issued and injected
→ redirect to APP_LOGIN_URL with {slug} filled in
```
Properties that must survive, verified by test:
- Nothing written to `instances` or `users` until the link is opened.
- `FindOneAndDelete` before provisioning, so a double-clicked link cannot create
two instances.
- Instance rollback if the owner insert fails, refusing to delete an instance
that has users.
- Re-submitting for the same address replaces the pending record.
- Rate limited to 3 signups per IP per hour, plus the honeypot field.
Two failure modes are new, because provisioning now spans two systems:
- **Paddle customer creation fails** — the instance and user are already created.
Complete the signup, record the account with an empty `PaddleCustomerID`, issue
the Free licence anyway, and flag for staff. A new customer must never be
blocked from signing in by a billing-system hiccup.
- **Licence issuance fails** — the instance exists with no licence and is
read-only. Flagged for staff, and the 15-minute reconciliation job (spec 3)
retries. The customer can log in and sees the licence banner.
Both resolve toward "the customer gets in", because a signup that half-fails
silently is worse than either outcome.
sitesvc changes: signup, verify, `site_pending_signups` and the provisioning
calls are deleted. `SITE_API_URL` gains a sibling for the admin endpoint, or the
marketing site posts signup to `ADMIN_API_URL` directly — the latter, so the two
form targets are explicit rather than implied.
### Configuration
| Variable | Required | Notes |
|---|---|---|
| `PADDLE_ENV` | yes | `sandbox` or `production`; selects which price IDs the plans table serves |
| `PADDLE_API_KEY` | yes | server-side API |
| `PADDLE_CLIENT_TOKEN` | yes | browser checkout; baked into the admin site build |
| `PADDLE_WEBHOOK_SECRET` | yes | signature verification. Boot fails without it — an unverified webhook endpoint is an endpoint anyone can issue licences through |
| `APP_LOGIN_URL` | yes | moved from sitesvc; `{slug}` template |
### Cutover
Signup migration is the only user-visible switch:
1. Deploy admin with signup enabled; sitesvc still serving its own.
2. Point the marketing site's form at admin. Deploy.
3. Let sitesvc's outstanding pending signups expire naturally — 24 hours — while
its verify endpoint stays live. **Do not delete the collection until it is
empty**, or someone's verification link breaks.
4. Deploy sitesvc with signup removed.
## Testing
**Webhooks:**
1. Each event type produces its documented action against a mock Paddle payload.
2. Replaying an event ID is a no-op returning `200`.
3. A bad signature is rejected `401` and processes nothing.
4. `subscription.updated` before `subscription.created` creates the subscription
and applies the update.
5. `subscription.canceled` issues nothing and leaves the current licence intact.
6. `past_due` leaves the licence intact and flags the account.
7. Renewal issues the next term, supersedes, resets `RelinkCount`, and the new
`ExpiresAt` is period end plus 3 days.
8. A handler failure writes to `admin_audit` and surfaces on the dashboard.
**Checkout:**
9. `custom_data` round-trips account, instance and tier through to the webhook.
10. Self-hosted checkout leaves the instance `awaiting_link` with no licence.
11. Linking after checkout issues the licence.
**Signup:**
12. Nothing is written to `instances` or `users` before the link is opened.
13. A double-clicked verification link creates exactly one instance.
14. Owner-insert failure rolls the instance back; rollback refuses an instance
with users.
15. Re-submitting replaces the pending record and invalidates the earlier link.
16. Rate limit and honeypot both reject.
17. Paddle customer creation failure still completes signup and issues the Free
licence.
18. Licence issuance failure still lets the user log in, showing the banner.
19. Expired pending records are dropped by the TTL index.
## Verification before merge
1. Full suite green.
2. Against Paddle **sandbox**, end to end for each tier: checkout with a test
card, confirm the licence is issued, confirm the instance reports `valid`.
3. Trigger a sandbox renewal and confirm the next term's licence arrives and is
injected.
4. Cancel in sandbox and confirm the licence keeps working to expiry, then the
instance degrades correctly — monitors still running.
5. Replay every webhook from Paddle's dashboard and confirm no duplicate licences
are created.
6. Full signup end to end through admin, then confirm the new user can log into
their control-plane instance and sees a valid Free licence.
7. Confirm sitesvc's pending-signup collection is empty before its signup code is
removed.
## Risks
| Risk | Mitigation |
|---|---|
| Duplicate licences from webhook retries | Unique index on event ID, checked before processing |
| Webhook missed entirely | 15-minute reconciliation job (spec 3) compares subscription state against issued licences |
| Cancellation not enforceable until expiry | Accepted, bounded by term; Self Hosted annual-only; stated in terms and on the cancellation screen |
| Signup cutover breaks in-flight verification links | Staged cutover; sitesvc's verify stays live until its collection is empty |
| Sandbox price IDs reaching production | `PADDLE_ENV` selects them from the plans table; environment badge in the admin site |
| Webhook endpoint unauthenticated | Signature verification mandatory; boot fails without the secret |
| Customer pays and never links | Reminder emails at 24h and 72h, staff dashboard alert at 48h |
@@ -1,286 +0,0 @@
# Spec 0a — Shared Module Extraction
Date: 2026-07-24
Status: Design approved, not implemented
Ships: independently. No dependency on any other licensing spec.
## Context
Vantage is three independent Go modules: `server`, `sitesvc`, `agent`. There is no
root `go.mod` and no `go.work`.
`sitesvc` writes into the same MongoDB collections the control plane reads, but
cannot import the control plane, so it carries hand-copied duplicates:
- `sitesvc/internal/models/models.go``Org` and `User` mirrored field for field
- `sitesvc/internal/provision/provision.go``Slugify`, `ReservedSlugs`,
`MinSlugLength`, `MaxSlugLength`, `BcryptCost`, slug-collision rules
Both files carry comments saying they must be changed in lockstep with the
control plane, and `CLAUDE.md` names the hazard explicitly: nothing enforces the
match. **The duplication has already drifted.** The control plane's `CreateOrg`
resolves slug collisions with an inline `fmt.Sprintf("%s-%d", base, i)` loop,
while sitesvc exposes the same rule as a separate `NextSlug(base, attempt)`
helper. They currently agree by luck, not by construction.
The licensing programme adds a fourth service (`admin`) that writes the license
blob onto the same tenant document. Adding a third copy of these rules is not
acceptable. This spec removes the duplication before any licensing code is
written.
This spec is a **pure refactor**. No database document changes. No behaviour
changes. Names stay as they are today (`Org`, `org_id`) — renaming happens in
spec 0b, deliberately kept separate so that a failed deploy has one suspect
rather than two.
## Goals
1. One authoritative definition of every document shape written by more than one
service.
2. One authoritative definition of provisioning rules (slug, bcrypt cost,
creation, rollback).
3. `sitesvc` keeps its independence from `server` — it depends on `shared`, not
on the control plane. The original design intent survives; only the copying
dies.
4. The agent is untouched.
## Non-goals
- Renaming anything. That is spec 0b.
- Moving control-plane-only models. `workflow.go`, `monitor.go`, `key.go`,
`server.go`, `secret.go`, `assignment.go`, `channel.go`, `console_session.go`,
`audit.go`, `org_oidc.go` stay in `server/internal/models`. Only the control
plane touches them, and hoisting them would make `shared` a dumping ground.
- Merging the repo into a single module.
## Design
### Module layout
```
vantage/
├── go.work # NEW: server, sitesvc, shared (NOT agent)
├── shared/ # NEW module: gitea.hostxtra.co.uk/mrhid6/vantage/shared
│ ├── go.mod
│ ├── models/
│ │ ├── org.go # Org
│ │ ├── user.go # User, RoleOwner/RoleAdmin/RoleMember, ValidRole
│ │ └── settings.go # Settings, AlertSettings, EmailSettings, SecretsSettings
│ ├── provision/
│ │ ├── slug.go # Slugify, BaseSlug, NextSlug, ReservedSlugs, limits
│ │ ├── org.go # CreateOrg
│ │ ├── user.go # CreateUser, BcryptCost
│ │ └── rollback.go # RollbackOrg
│ └── indexes/
│ └── indexes.go # EnsureCoreIndexes
├── server/ # replace => ../shared
├── sitesvc/ # replace => ../shared
└── agent/ # untouched
```
`go.work`:
```
go 1.26
use (
./shared
./server
./sitesvc
)
```
Each consumer's `go.mod` also carries an explicit replace:
```
require gitea.hostxtra.co.uk/mrhid6/vantage/shared v0.0.0
replace gitea.hostxtra.co.uk/mrhid6/vantage/shared => ../shared
```
Both are needed. `go.work` makes editors, `go test ./...` and local tooling work
across modules. The `replace` directives make Docker builds work whether or not
`go.work` is present, and stop `go build` outside the workspace from silently
trying to resolve `shared` from the network.
`shared` depends only on `go.mongodb.org/mongo-driver/v2`,
`golang.org/x/crypto/bcrypt` and `github.com/google/uuid`. It must not import
gin, redis, guac or anything else from the control plane's tree — that is what
keeps sitesvc small.
### What moves
**`shared/models`** — the three documents written by more than one service:
| Type | From | Written by |
| ------------------------------------- | ------------------------------------ | ---------------------------------- |
| `Org` | `server/internal/models/org.go` | server, sitesvc, later admin |
| `User` + role constants + `ValidRole` | `server/internal/models/user.go` | server, sitesvc |
| `Settings` and its sub-structs | `server/internal/models/settings.go` | server today; admin reads it later |
`Settings` moves now rather than later because spec 3's admin service reads it,
and moving it later would mean a second round of import churn across both
services.
`PendingSignup` does **not** move. Only sitesvc writes `site_pending_signups`,
and the control plane does not know the collection exists.
**`shared/provision`** — the rules, promoted from private helpers to a real API:
```go
const (
MinSlugLength = 3
MaxSlugLength = 40
BcryptCost = 12
)
var ReservedSlugs = map[string]bool{ /* www, api, app, admin, auth, install, static, _next, default */ }
func Slugify(name string) string
func BaseSlug(name string) (string, error) // validates length + reserved
func NextSlug(base string, attempt int) string
// CreateOrg resolves a free slug and inserts. The caller supplies the
// collection handle so shared does not own a Mongo connection.
func CreateOrg(ctx context.Context, db *mongo.Database, name string) (*models.Org, error)
func CreateUser(ctx context.Context, db *mongo.Database, orgID, email, password, role string) (*models.User, error)
// RollbackOrg deletes an org only if it has no users. Refuses otherwise.
func RollbackOrg(ctx context.Context, db *mongo.Database, orgID string) error
```
`shared.CreateOrg` becomes the single implementation. The control plane's
`services.CreateOrg` shrinks to a wrapper that calls it and then runs
`SeedDefaultSteps` — seeding stays in the server, because `shared` must not know
about workflow steps. sitesvc calls `shared.CreateOrg` directly and does not
seed, which is the behaviour it has today.
Note on the slug loop: it is count-then-insert and therefore racy. It is safe
only because of the unique index on `orgs.slug`. `CreateOrg` must keep handling
`mongo.IsDuplicateKeyError` and returning a clean error — moving the code must
not lose that. Document the reliance in a comment at the loop.
**`shared/indexes`** — `EnsureCoreIndexes(ctx, db)` declares the unique indexes
on `users.email` and `orgs.slug`. Both services call it at boot; creating an
existing index is a no-op. These indexes are a security property, not an
optimisation (see `CLAUDE.md`: `GetUserByEmail` does an unscoped `FindOne`, so
duplicates would break the OIDC cross-org guard), so the shared version is
**fatal on failure** for both callers.
Server-only index builders (`EnsureSettingsIndexes`, `EnsureSecretIndexes`,
`EnsureWorkflowIndexes`) stay in the server and keep their current
fatal/warn behaviour.
### What is deleted
- `sitesvc/internal/models/models.go` — reduced to `PendingSignup` only
- `sitesvc/internal/provision/` — deleted entirely
- `sitesvc/internal/store/store.go` — org/user creation replaced by calls into
`shared/provision`; pending-signup storage stays
### Docker and CI
Both Go Dockerfiles currently build with the module directory as context:
```dockerfile
WORKDIR /app
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN go build ... ./cmd
```
A `replace => ../shared` cannot resolve from that context. Build contexts move
to the repo root:
```dockerfile
WORKDIR /src
COPY shared/go.mod shared/go.sum ./shared/
COPY server/go.mod server/go.sum ./server/
RUN cd server && go mod download
COPY shared/ ./shared/
COPY server/ ./server/
ARG VERSION=dev
RUN cd server && CGO_ENABLED=0 GOOS=linux go build \
-ldflags="-s -w -X main.Version=${VERSION}" -o /vantage-server ./cmd
```
The two-stage copy keeps the dependency-download layer cached, which is the
reason the current Dockerfiles are written the way they are.
`.gitea/workflows/server-deploy.yml` must set `context: .` and
`file: server/Dockerfile` (and likewise for sitesvc) for the two Go images. The
`web` and `site` image builds are unaffected.
`agent-release.yml` is untouched. The agent is not in the workspace, has no
`replace`, and cross-compiles exactly as it does today.
### Error handling
No new error paths. `shared/provision` returns the same error strings the two
callers produce today so that API responses do not change. The one place to be
careful is wording: sitesvc says "organisation" and the control plane says
"organization". `shared` standardises on **"organisation"**; the control plane's
two error strings change spelling. This is user-visible in API error text and is
called out here so it is a decision rather than an accident.
## Testing
**No automated tests.** Decision taken 2026-07-24: the repo has no Go test suite
and one is not being started here. Verification is by compiler, `grep`, and
running both services end to end.
That places the whole weight on three manual checks, which the implementation
plan makes mandatory steps rather than suggestions:
1. **bson tag diff**`diff` the `bson:"…"` tags of each moved struct against
the originals. A changed tag orphans production data silently, and this is
the only thing that catches it.
2. **Slug behaviour walkthrough** — a throwaway `main` printing `Slugify`,
`BaseSlug` and `NextSlug` output for a fixed input table, compared against
expected output recorded in the plan.
3. **End-to-end agreement** — sign up through sitesvc against a scratch
database, open the verification link, then log into the control plane with
those credentials. This is the check that proves the two services still agree
about the documents they share. If it passes, the refactor worked.
Plus `grep` assertions that exactly one definition of `Slugify` and
`ReservedSlugs` survives repo-wide, and that no struct under `sitesvc/` carries
a `bson:"org_id"` tag.
## Verification before merge
Evidence required, not assertions:
1. `go build ./...` succeeds in `shared`, `server` and `sitesvc`.
2. `go vet ./...` clean in all three.
3. `docker build -f server/Dockerfile .` and `docker build -f sitesvc/Dockerfile .`
both succeed from the repo root.
4. `grep -r "org_id" sitesvc/` returns hits only in `PendingSignup` context and
`shared` imports — no local struct redefinitions.
5. End-to-end against a scratch database: sitesvc signup form → verification link
→ org and owner created → that owner logs into the control plane
successfully. This is the test that proves the two services still agree.
6. The agent still builds for `linux/amd64`, `linux/arm64` and `windows/amd64`.
## Rollout
Single release. `server` and `sitesvc` images must be deployed together — a skew
is harmless here (documents are unchanged) but there is no reason to split it.
No database migration. No downtime.
## Risks
| Risk | Mitigation |
| -------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
| Docker context change breaks CI | Verified locally by building both images from root before pushing |
| Behaviour drift while moving `CreateOrg` | Unit tests written against current behaviour first, then the move |
| `shared` accumulating control-plane concerns | Explicit non-goals above; keep its `go.mod` dependency list to three entries and review any addition |
| Error-string spelling change | Called out as a decision; grep the web UI for hard-coded matches on the old strings |
## Follow-on
Spec 0b (`instance-rename`) becomes a rename inside one module plus its
consumers, rather than a rename across three independent copies. That is the
whole reason this spec goes first.
@@ -1,498 +0,0 @@
# Cloud instance creation and account membership — design
Spec 6. Designed 2026-07-26. Depends on specs 0a, 0b, 1, 2 and 3, all shipped.
## The problem
`https://vantage.hostxtra.co.uk/start` provisions a control-plane instance the
moment a customer opens the verification email. It creates nothing on the admin
side: no `accounts` row, no `admin_instances` row, no licence. Every cloud
customer who signs up today therefore lands on an unlicensed instance that spec 2
degrades to read-only, and staff must attach it by hand afterwards.
The fix is not to bolt licence issuance onto the existing verification handler.
It is to separate the two things that flow has conflated — **having an account**
and **having an instance** — so that the account exists first and the instance is
something the customer asks for.
Once accounts are real, a second thing follows: an account has *people* in it,
and those people need access to the account's instances. That is what forces the
`users.email` change below, and it is the largest single item in this spec.
## The new flow
```
site/start form ──POST──▶ admin /auth/signup
account name, email, password
→ accounts row + unverified customer_users row (account_role owner)
→ verification email; nothing written to the control plane
verification link ──▶ admin /auth/verify
→ customer_users.verified_at set
→ customer signs in at vantage-hq.hostxtra.co.uk
HQ portal, "Create instance" ──POST──▶ admin /api/instances
→ control-plane instances + users (creator becomes owner)
→ admin_instances row + instance_members row
→ Free licence issued, injected, emailed as "your instance is ready"
HQ portal, "Invite" and "Add to instance"
→ more customer_users on the account
→ each grant projects a control-plane users row into that cloud instance
```
Signup itself needs almost no new code: `auth.HandleSignup` already creates an
account, an unverified `customer_users` row and a verification email, and was
written for self-hosted customers. It turns out to be exactly the account-first
signup cloud needs.
This supersedes the "Signup migration off sitesvc" section of spec 5
(`2026-07-24-paddle-billing-design.md`). That section moved the *existing*
signup-provisions-an-instance flow to admin unchanged; this changes its shape.
Spec 5's Paddle work is unaffected and layers on top: the £0 Free subscription
and the Paddle customer are created where this spec issues the Free licence.
Paddle is explicitly **out of scope here**. Accounts created by this spec have an
empty `PaddleCustomerID`, which spec 5's account model already permits.
## Phasing
Three phases, each shippable, in this order. The plan should not interleave them
— phase 1 changes an index that everything else then depends on.
1. **Identity** — drop the global email index, scope the two unscoped lookups,
add the `hq` fields to the user document, and remove admin's unscoped
control-plane login branch. No new UI, and nothing is projected yet.
2. **Instance creation and Free lifecycle**`POST /api/instances`, renewal,
notices, the reaper, the sitesvc cutover.
3. **Membership** — account roles, invitations, per-instance grants, password
propagation.
## Phase 1 — identity
### Dropping the global email index
`users.email` currently carries a unique index **across the whole control
plane**. `CLAUDE.md` names it a security property, and it is one today. It is
also what makes "an account's people belong to several instances" impossible:
one address can own exactly one user document anywhere.
It is replaced by a unique compound index on `(instance_id, email)`, which is the
constraint that was actually wanted: one address is one user *within an
instance*.
The global index is only load-bearing because two lookups are unscoped. Both are
scoped instead, and the scoped lookups are a strictly stronger guarantee than the
index was — an index prevents the ambiguity, whereas a scoped query cannot be
ambiguous in the first place.
| Caller | Today | After |
|---|---|---|
| `auth.HandleLocalLogin` | `GetUserByEmail(email)` | resolve the instance, then `GetUserInInstanceByEmail` |
| `auth.HandleOIDCCallback` | `GetUserByEmail(email)`, then a cross-instance guard | `GetUserInInstanceByEmail(instanceID, …)`; the guard is deleted as unreachable |
`services.GetUserByEmail` is **deleted**, not merely left unused. Leaving an
unscoped helper in place is how this bug comes back.
Resolving the instance for local login:
1. `InstanceFromHost` — always succeeds on cloud, where every instance has its
own subdomain.
2. Otherwise, if exactly one instance exists, use it. This is the self-hosted
case, which is single-instance by construction because a licence binds one
instance UUID.
3. Otherwise refuse with a message naming the cause, rather than guessing.
### The index migration
In `shared/indexes.EnsureCoreIndexes`, in this order:
1. Create the unique compound index on `(instance_id, email)`. Fatal on failure.
2. Drop `email_1` if present, ignoring `IndexNotFound` so it is idempotent.
Creating before dropping means a failure at step 2 leaves both indexes in place,
which is safe. A failure at step 1 leaves the old index alone, which is also
safe.
`EnsureCoreIndexes` is called at boot by server, sitesvc and admin, so **all
three images must ship together**. `server-deploy.yml` rebuilds every image on
every push to `main`, so this happens by default; the risk is only a partial
manual rollout on the host.
**This migration is one-way.** Once two users share an address across instances,
`email_1` cannot be recreated. Rolling the server back past this change would
leave unscoped lookups running against data that can now be ambiguous. The
rollback plan is forward-only: fix and redeploy.
`sitesvc.EmailTaken` also does an unscoped count over `users`. It disappears with
sitesvc's signup in phase 2.
`admin.HandleCloudLogin`'s control-plane branch does an unscoped
`FindOne({email})` too, and unlike the other two there is no instance in context
to scope it by — HQ login is not per-instance. **That branch is deleted.** Every
customer created by this spec has a `customer_users` row, which already wins in
the existing precedence. Legacy cloud customers are handled by staff, who already
attach their instances by hand per the spec README, and who gain
`POST /api/staff/accounts/:id/users` to create their HQ login.
### The control-plane user document
`shared/models.User` gains:
- `hq_user_id` — the `customer_users.user_id` this row was projected from, absent
on locally-created users.
- `auth_source: "hq"` as a third value alongside `local` and `oidc`.
An `hq`-sourced user is **managed in HQ, not in the instance**. The control plane
refuses to change its role, delete it, or change its password through
`/api/instance/users`, answering with "managed in Vantage HQ". `web/` renders
those rows read-only with the same label. Locally-created users are unaffected
and stay fully editable in the instance — a cloud instance can hold both kinds.
This gives one owner per fact. A role that is editable in two places is a role
with two answers.
## Phase 2 — instance creation
`POST /api/instances`, customer session, `account_role` owner or admin,
body `{ "name": "..." }`.
In order, each step undoing the previous on failure:
1. Refuse if the account already holds a non-cancelled Free instance — a
pre-check of the same rule `licensing.checkFreeLimit` enforces, so we never
create an instance we then cannot licence. `409`.
2. Read the caller's `customer_users` row for its bcrypt hash.
3. `provision.CreateInstance` — control-plane instance and slug.
4. `provision.CreateUserWithHash(…, RoleOwner, "hq")` with that hash and
`hq_user_id`. On failure, `provision.RollbackInstance`.
5. Insert `admin_instances`, then `instance_members` for the creator. On failure,
delete the control-plane user, then roll back the instance.
6. `licensing.Issue{Tier: free, Term: "monthly", Reason: ReasonNew, IssuedBy:
"self-serve"}`, then `inject.Deliver`.
7. Email the creator: instance URL, sign-in address, licence expiry date.
Steps 6 and 7 do **not** fail the request. A licence that was not issued is
recoverable — the instance exists, the customer can sign in, they see spec 2's
licence banner, and staff can issue by hand. Failing the whole creation and
rolling back an instance the customer can already see would be worse. This
matches the rule spec 5 states for the same pair of failures: both outcomes
resolve toward "the customer gets in".
### Admin's control-plane write boundary
`inject`'s package doc says plainly that a second write target into the control
plane "is a design change and not a refactor". This is that design change, and it
is made explicitly rather than by widening `inject`.
Provisioning and membership projection live in a **new package,
`admin/internal/cloudprov`**. `inject` is left untouched, still writing exactly
three licence fields on `instances`. `db` gains a `ControlDB() *mongo.Database`
accessor, because `shared/provision` takes a database rather than a collection.
`cloudprov` writes exactly three things: instance documents (create and roll
back), user documents (create, delete, update role and password hash), and
nothing else. `CLAUDE.md`'s description of the boundary is updated in the same
commit, because it currently claims admin's control-plane access is read-only
apart from three licence fields, and that stops being true here.
## Phase 2 — Free lifecycle
A Free licence runs for one month plus the existing three-day `GracePeriod`,
using the `"monthly"` term `licensing.Issue` already implements. No new term
value. One Free instance per account, unchanged.
### Renewal
`POST /api/instances/:id/renew`, through `ownedInstance`, account owner or admin.
- Tier must be Free. Paid tiers renew through billing, not here.
- Allowed once `now > expires_at - 7d`, and at any point after that up to
deletion — so the same button rescues a lapsed instance rather than needing a
second mechanism.
- Reissues Free with `Reason: ReasonRenewal`, injects, emails the new date.
Renewal is deliberately manual. It is the entire reclaim signal: an instance
nobody renews is an instance nobody is using.
### Status and notices
`admin_instances.status` gains `deleted`. A sweep in admin flips `active` to
`lapsed` when the current licence's `expires_at` passes, and the existing
15-minute reconciler — which already logs "no control-plane instance X" — flips
those to `deleted` and clears their `instance_members` rows instead of only
logging.
Four emails to the account's owners and admins, driven by `expires_at`:
| When | Says |
|---|---|
| 7 days before expiry | Renew, one click, here is the link |
| on expiry | Read-only now; deleted in 14 days unless renewed |
| 7 days before deletion | Deleted in 7 days |
| 1 day before deletion | Deleted tomorrow |
Each send is recorded on the `admin_instances` document, so a restart or a double
tick cannot re-send one. Renewal clears the record, so the next term starts the
sequence again.
## Phase 2 — deletion
Deletion is the only irreversible path in the system, so it is owned by the
service that knows what an instance is made of.
**The reaper runs in the control plane, not in admin.** Admin already injects
`license_tier` and `license_expiry` onto the instance document, so the server
drives off data it holds locally, and the list of collections carrying
`instance_id` stays in the codebase that defines them. Mirroring that list into
admin would be exactly the class of duplication `CLAUDE.md` already warns about
for slug rules and design tokens — except a divergence here deletes the wrong
rows or leaves orphans behind.
The sweep, in `server/internal/services`:
- Eligible when `license_tier == "free"` **and** `license_expiry` is present
**and** `license_expiry` is more than the configured window in the past.
- Purges the instance document, its users, and every `instance_id`-scoped
document across the collections listed in `CLAUDE.md`. Workflow run logs on
disk go with them.
- Fail-safe by construction. An instance whose licence issuance failed has no
`license_tier` and is never eligible. A paid instance is never eligible. An
instance admin has not reached yet keeps whatever expiry was last injected, and
admin's reconciler keeps that field current.
- Every purge writes an audit entry before deleting, and logs the instance ID,
slug and document counts.
- Admin's reconciler notices the instance has gone and cleans up its own
`admin_instances` status and `instance_members` rows.
### The kill switch
Gated on `FREE_INSTANCE_REAP_AFTER`, a duration. **Empty disables the sweep
entirely**, and empty is the default.
It is unset in `deploy/docker-compose.yml` and set to `336h` only in
`deploy/docker-compose.site.yml`, so a self-hosted deployment can never reap
anything — the same containment rule that keeps `LICENSE_SIGNING_KEY` in exactly
one service in exactly one compose file.
## Phase 3 — accounts, people and membership
### The model
```
Account
├── customer_users the people. account_role: owner | admin | member
└── admin_instances the deployments
└── instance_members which people are on which cloud instance
```
`customer_users` gains `account_role`. Existing rows backfill to `owner` — they
are all account creators today. Owners and admins may invite users, create
instances, and grant instance access; billing stays owner-only. The vocabulary
deliberately matches the control plane's own three roles rather than inventing a
second one.
`instance_members` is new: `{member_id, account_id, instance_id,
customer_user_id, role, control_user_id, created_at}`, unique on
`(instance_id, customer_user_id)`. `role` is the role the projected
control-plane user holds inside the instance.
### Grants project, they do not federate
Granting a user access to a cloud instance creates a real control-plane `users`
row through `cloudprov`, with `auth_source: "hq"` and `hq_user_id` set. The
instance authenticates it exactly as it authenticates any other user, with no
runtime dependency on admin. Revoking deletes that row.
**Self-hosted instances are never projected into.** `POST /api/instances/:id/
members` refuses when `deployment != cloud`, with that as the message. For a
self-hosted instance the account's users exist to manage the licence, and the
instance's own users are managed locally in the customer's own deployment, which
we cannot see and have no business writing to.
Endpoints, all customer-session and all through `ownedInstance` where an instance
is named:
```
GET,POST /api/account/users invite; owner|admin
PUT /api/account/users/:id/role owner|admin; cannot demote the last owner
DELETE /api/account/users/:id owner|admin; revokes every grant first
PUT /api/account/password any user; propagates
GET,POST /api/instances/:id/members owner|admin
PUT /api/instances/:id/members/:uid/role
DELETE /api/instances/:id/members/:uid
```
Invitations reuse `auth.CreateCustomerUser`, which already does the
unverified-row-plus-verification-email dance and already deletes the row if the
email fails to send. A user cannot be granted an instance until verified.
Revoking the last **owner** of an instance is refused, mirroring the control
plane's own `ErrLastOwner`. The check counts control-plane owners for that
instance, so it also sees owners created locally inside the instance.
### Password propagation
The HQ password is the single source of truth for every `hq`-sourced row.
`PUT /api/account/password` rehashes at cost 12, updates `customer_users`, then
has `cloudprov` write the same hash to every control-plane user carrying that
`hq_user_id`. The instance refuses to change an `hq`-sourced user's password
locally, so there is no competing writer.
Propagation is best-effort and retried, on exactly the pattern `inject` already
proves: a failure is logged and flagged, and admin's 15-minute reconciler gains a
pass that compares each `hq`-sourced row's hash against its `customer_users`
source and repairs mismatches. The worst case is a stale password on one instance
for up to fifteen minutes, which is recoverable; failing the password change
because one of three instances was unreachable is not.
## Frontend
### `site/`
`components/InstanceForm.tsx` becomes `AccountForm.tsx`: account name, email,
password, honeypot. It posts to `NEXT_PUBLIC_ADMIN_API_URL/auth/signup` rather
than to sitesvc. The live `your-instance.vantage.hostxtra.co.uk` slug preview
goes — there is no instance yet at this point, and showing one would be a lie.
`app/start/page.tsx` copy changes from "Set up your instance" to creating an
account, and its "What happens next" panel gains the create-an-instance step
between confirming the email and adding a key.
`ADMIN_API_URL` gains a browser-reachable presence in the `site` image build, and
`site`'s origin must be listed in admin's `ADMIN_ORIGIN`. Both are new failure
modes with the same footgun `CLAUDE.md` already documents for `SITE_API_URL`.
`SITE_API_URL` still serves the contact form.
### `adminsite/`
- `(customer)/page.tsx` — the "No instances yet" panel gains a primary **Create a
free instance** action. Hidden once the account holds a Free instance, with the
reason stated rather than the button silently absent.
- `(customer)/instances/new/` — name field and a live slug preview of the
resulting `<slug>.vantage.hostxtra.co.uk`.
- `(customer)/instances/[id]/` — a members panel: who is on this instance, their
role, add and remove. Absent for self-hosted instances, replaced by a line
saying users are managed inside the install.
- `(customer)/users/` — the account's people, invitations, account roles.
- `(customer)/settings/` — change password, with a note that it applies to every
instance you belong to.
- `components/InstanceCard.tsx` — expiry date, a **Renew** action inside the
window, and a deletion countdown when lapsed. Per `CLAUDE.md`'s rule, licence
state never reads by colour alone; the countdown is a text label.
- `lib/api.ts` — the new calls, `"deleted"` on `InstanceStatus`, and an
`AccountRole` type.
### `web/`
`settings/instance` gains the read-only treatment for `hq`-sourced users: role
shown, controls disabled, labelled "managed in Vantage HQ" with a link to the
portal. Everything else is unchanged; spec 2's licence banner already covers a
lapsed instance.
## sitesvc
Signup, verify, `site_pending_signups`, `EmailTaken` and the provisioning calls
are deleted. sitesvc keeps the contact form only, and drops `APP_LOGIN_URL`.
The staged cutover from spec 5 applies unchanged, and matters for the same
reason: an in-flight verification link must not break.
1. Deploy admin. Its signup already exists; nothing to enable.
2. Point `site/start` at admin. Deploy `site`.
3. Wait for sitesvc's outstanding pending signups to expire — 24 hours — with its
verify endpoint still live. **Do not delete the collection until it is empty.**
4. Deploy sitesvc with signup and verify removed.
A signup that completes through the old path during step 3 produces an instance
with no account and no licence, exactly as today. Staff attach those by hand, the
same job the README already describes for existing cloud tenants.
## Configuration
| Service | Variable | Required | Notes |
|---|---|---|---|
| admin | `APP_LOGIN_URL` | yes | moved from sitesvc; `{slug}` template, used in the instance-ready email |
| server | `FREE_INSTANCE_REAP_AFTER` | no | duration past expiry before a Free instance is purged. **Empty disables the reaper**, and empty is the default. `336h` in `docker-compose.site.yml` only |
| site build | `ADMIN_API_URL` | yes | browser-reachable; must be in admin's `ADMIN_ORIGIN` |
| sitesvc | `APP_LOGIN_URL` | — | removed |
## Testing
Phase 1, identity:
1. The compound index exists and `email_1` is gone after one boot; a second boot
is a no-op.
2. Two users with the same address in different instances can both be created and
both sign in, each landing in their own instance.
3. Two users with the same address in one instance are refused by the index.
4. Local login on a cloud subdomain finds only that instance's user; the same
address on another instance is not reachable from this host.
5. Local login on a bare host with one instance works; with two it refuses with a
named cause rather than picking one.
6. OIDC provisions into the instance from the callback state, and an address
belonging to another instance no longer produces a cross-org error because it
is simply not found — it provisions a new member instead, which is correct.
7. `GetUserByEmail` no longer exists.
Phase 2, creation and lifecycle:
8. Signup writes nothing to `instances` or `users`; only the emailed link makes
the account usable.
9. Creating an instance produces an instance, an `hq`-sourced owner user, an
`admin_instances` row, an `instance_members` row, a Free licence, and an
injected `license_blob`.
10. The creator can sign in to the new instance with their HQ password.
11. A second Free instance on the same account is refused `409` and writes
nothing.
12. Owner-insert failure rolls the instance back, and rollback refuses an
instance that has users.
13. Licence issuance failure still leaves a signed-in-able instance and flags for
staff.
14. Renew outside the window is refused; inside it, it supersedes, injects and
moves `expires_at` forward by a month plus grace.
15. Renewing a lapsed instance restores it before the reaper takes it.
16. Each notice sends once across a restart.
Phase 2, the reaper — the part that must be got right:
17. With `FREE_INSTANCE_REAP_AFTER` empty, nothing is ever deleted.
18. An instance with no `license_tier` is never eligible, whatever its age.
19. A Professional instance past expiry is never eligible.
20. A Free instance one hour short of the window is not deleted; one hour past it
is.
21. A purge leaves no document carrying that `instance_id` in any collection, and
writes an audit entry first.
22. Purging is idempotent — a second run over a half-deleted instance completes
it rather than erroring.
Phase 3, membership:
23. An invited user cannot be granted an instance until verified.
24. A grant creates a control-plane user that can sign in to that instance with
the invitee's HQ password.
25. The same user can hold rows in two instances at once, with different roles.
26. Revoking deletes the control-plane row, and that user can no longer sign in
to that instance while keeping access to the others.
27. Revoking or demoting an instance's last owner is refused, including when that
owner was created locally inside the instance.
28. Granting against a self-hosted instance is refused and writes nothing to the
customer's deployment.
29. A `member` cannot invite, create instances, or grant access.
30. A password change propagates to every linked instance; with one instance's
write forced to fail, the reconciler repairs it within one pass.
31. An `hq`-sourced user's role, deletion and password are refused inside the
instance API, not merely hidden in `web/`.
## Risks
| Risk | Mitigation |
|---|---|
| Dropping `email_1` is one-way and weakens a documented security property | Scoped lookups ship in the same binary that drops the index; the unscoped helper is deleted so it cannot be reintroduced; the compound index restores the equivalent guarantee; rollback plan is forward-only and stated |
| A partial rollout leaves an old service recreating `email_1` | All three services call `EnsureCoreIndexes`; `server-deploy.yml` rebuilds every image per push; the host rollout command already updates all services together |
| Reaper deletes a live instance | Kill switch defaults off; eligibility needs an explicitly-Free tier and a present expiry; unset fields are never eligible; four warning emails precede it |
| Admin's widened control-plane write access grows further | Confined to `cloudprov`, which writes instances and users and nothing else; `inject` untouched; `CLAUDE.md` updated to say so |
| Password propagation leaves an instance stale | Reconciler pass compares and repairs; worst case is fifteen minutes; the instance refuses local changes so there is no competing writer |
| A projected user is edited in both places | `hq`-sourced rows are refused by the instance API, not merely hidden in the UI |
| Cutover breaks an in-flight verification link | sitesvc's verify stays live until its collection is empty |
@@ -1,455 +0,0 @@
# Metered Licensing — Design
**Status:** designed 2026-07-26. Supersedes parts of spec 5 (paddle-billing) and
the tier table in [`README.md`](README.md).
**Goal:** turn the licence from a snapshot of a fixed tier into a snapshot of
what one customer configured and paid for. Two deployments times three tiers,
servers metered per month, features opted into individually, all of it
self-service in Vantage HQ.
**Why now:** spec 5 is designed but not implemented — `admin/internal/paddle`
and `admin/internal/billing` do not exist. Its `Subscription` struct, its
`plans.paddle_price_ids` shape, its single-price checkout and its
`ApplySubscription` all assume one price per subscription, and a metered plan has
several. Folding this in now costs a revision of an unstarted plan; folding it in
later would cost a rewrite of shipped billing code.
---
## The pricing model
Two deployments, three tiers, six plans.
| | servers | monitors | secret groups | channels | audit history | console | SSO | support |
|---|---|---|---|---|---|---|---|---|
| **Free** | 3 | 3 | 1 | 1 | 30 days | — | — | community |
| **Professional** | 3 + N | ∞ | ∞ | ∞ | 365 days | opt-in | opt-in | email, 24/5 |
| **Enterprise** | 10 + N | ∞ | ∞ | ∞ | ∞ | opt-in | opt-in | email + call, 24/7 |
The allowances are identical in both deployments. What differs is the term:
| | monthly | annual |
|---|---|---|
| Cloud Free | — | yes, renewed from HQ |
| Cloud Professional | yes | yes |
| Cloud Enterprise | yes | yes |
| Self-Hosted Free | — | yes, renewed from HQ |
| Self-Hosted Professional | — | yes |
| Self-Hosted Enterprise | — | yes |
**Self-Hosted stays annual-only, for the reason already written into
`shared/license/license.go`:** an offline licence cannot be revoked, so the term
length *is* the revocation window. A self-hosted monthly licence would renew that
unrevokable window twelve times a year for no commercial gain. A resolved
self-hosted monthly price is therefore a configuration error and must fail loudly
rather than issue.
**Servers are the only metered dimension.** Everything above Free is unlimited
except audit history. This was a deliberate narrowing: an earlier draft sold
secret groups in blocks of five, and dropping it leaves one number for a customer
to understand and one line item on an invoice.
**Enterprise is self-service at a published price**, bought through the same
configurator as Professional. The 24/7 phone commitment is an operational promise
we make, not a technical gate we build.
**Support level is not enforced by anything.** It is carried for display, and
that is the whole of its job.
---
## What breaks, and must be fixed in the same change
Three invariants stop being true. Each is load-bearing today.
**`plans` is keyed on `tier` alone.** It becomes `(deployment, tier)` with a
unique index on the pair. `license.PlanFor(tier)` becomes
`PlanFor(deployment, tier)`.
**Free is cloud-only by construction.** The single comparison in
`licensing.Issue``plan.Deployment != inst.Deployment` — is what enforces it
today, because Free's only plan row says `cloud`. With a self-hosted Free row
that comparison stops meaning "Free is cloud-only" and starts meaning only "the
plan row matches the instance". The paragraph in `shared/license/plans.go`
claiming construction-level enforcement must go, because it is no longer true.
**`checkFreeLimit` counts Free instances per account.** It must count per account
*and deployment*, or a customer holding a cloud Free instance is refused a
self-hosted Free one with a message about a limit they have not reached.
---
## Data model
### `plans` — the tier definition
Loses `paddle_product_id` and `paddle_price_ids` entirely; those move to
`catalogue`. Safe to delete because nothing has ever written to them.
```
{deployment: "cloud", tier: "professional", name: "Professional",
base_limits: {max_servers: 3, max_monitors: -1, max_secret_groups: -1,
max_channels: -1, audit_retention_days: 365},
base_features: [], support_level: "email_24_5", active: true}
```
`base_limits` replaces `limits`: it is the allowance before anything is bought,
which is a different claim from the one the old field made. `base_features` is
what the tier includes without opting in — empty for all six plans today, because
console and SSO are both opt-in, but the field is what lets a future tier bundle
one.
### `catalogue` — every priceable component
The only place a Paddle price ID appears anywhere in the system.
```
{kind: "base", deployment: "cloud", tier: "professional",
price_ids: {sandbox: {monthly: "pri_…", annual: "pri_…"},
production: {monthly: "pri_…", annual: "pri_…"}}}
{kind: "limit", deployment: "cloud", tier: "professional", limit_key: "max_servers",
price_ids: {sandbox: {monthly: "pri_…", annual: "pri_…"}, production: {…}}}
{kind: "feature", deployment: "cloud", tier: "professional", feature_key: "console",
price_ids: {}}
{kind: "feature", deployment: "cloud", tier: "professional", feature_key: "oidc",
price_ids: {}}
```
Unique index on `(deployment, tier, kind, limit_key, feature_key)`.
- **`kind: "base"`** is the plan's own fee, always quantity 1.
- **`kind: "limit"`** raises a named limit by one per quantity. `limit_key` is a
field name in `license.Limits`, so adding metered channels later is a catalogue
row and no code. There is deliberately **no `block_size` field**: with
secret-group blocks dropped it would be `1` in every row that will ever exist.
- **`kind: "feature"`** is a feature key. **An empty `price_ids` means free to
toggle.** A price appearing later is a staff edit in the plans UI, not a
migration and not a deploy — which is the whole reason features are catalogue
rows rather than a list on the plan.
A self-hosted row simply has no `monthly` key. Nesting by environment before term
keeps promoting sandbox to production a configuration change, as spec 5 already
established.
### `entitlements` — one row per instance
The customer's configuration. Both the subscription and the licence are derived
from it; it is derived from nothing.
```
{instance_id: "uuid", account_id: "uuid",
deployment: "cloud", tier: "professional", term: "monthly",
desired: {servers: 10, features: ["console"]},
granted: {servers: 5, features: []},
resolved_limits: {max_servers: 5, max_monitors: -1, max_secret_groups: -1,
max_channels: -1, audit_retention_days: 365},
granted_at, updated_at, scheduled_change_at}
```
Unique index on `instance_id`.
**`desired` is what they asked for; `granted` is what a payment confirmed.** The
checkout and the subscription update are built from `desired`. A licence is only
ever signed from `granted`. An abandoned checkout therefore leaves a `desired`
that reached no licence, which is harmless, and HQ can say "pending change"
truthfully instead of guessing.
**`resolved_limits` is stored, not derived on read.** It is `plan.base_limits`
with `granted.servers` folded in, and it is what `Issue` snapshots. Storing it
keeps the fold in exactly one place; deriving it at every read would put the
arithmetic in the issuer, the portal and the staff console.
**Free gets a row at instance creation** with `desired == granted` and no
subscription. Every one of the six cases then reads the same shape, and licence
issuance has one path rather than a Free branch.
### `license.Limits` gains two fields
```go
type Limits struct {
MaxServers int `json:"max_servers"`
MaxMonitors int `json:"max_monitors"`
MaxSecretGroups int `json:"max_secret_groups"`
MaxChannels int `json:"max_channels"`
AuditRetentionDays int `json:"audit_retention_days"`
}
```
`MaxMonitors` behaves exactly like the existing counts. `AuditRetentionDays` is a
new kind of limit — a duration rather than a cap — and `Unlimited` means never
trim.
### `license.License` gains `SupportLevel string`
Display-only, exactly as `InstanceName` already is. It goes in the signed payload
rather than being fetched from HQ so that `/settings/license` can state the
support level on an air-gapped install, which is the one deployment most likely
to need to know who to call.
### `models` additions
`ReasonEntitlementChange = "entitlement_change"` joins the issuance reasons.
Reasons end up in support conversations, so a mid-term server addition must not
be filed as a renewal — a renewal resets `relink_count`, and adding a server is
not a new term.
---
## Resolution
Two folds, in one package (`admin/internal/catalogue`), so the arithmetic exists
once.
**To a licence.** `Resolve(plan, granted) → (license.Limits, []string)`:
start from `plan.base_limits`, and for each `kind: "limit"` row add the
configured quantity to `limit_key`. `granted.servers` is the *total* the customer
sees, so the quantity billed is `servers - plan.base_limits.max_servers` and the
resolved limit is `servers`. Features are `plan.base_features` plus
`granted.features`, deduplicated, filtered to keys the catalogue actually offers
for that `(deployment, tier)` — a stale feature key in a stored entitlement must
not survive into a signed payload.
**To Paddle line items.** `LineItems(env, deployment, tier, term, desired) → []Item`:
the base row at quantity 1, the server row at quantity
`desired.servers - base_limits.max_servers`, and one item per desired feature
that has a price ID in this environment and term. A feature with no price ID
produces no line item and is granted for free. A quantity of zero produces no
line item at all, so a Professional customer at exactly 3 servers has a
single-item subscription.
**Reverse resolution replaces spec 5's `ResolvePriceID`.** A metered subscription
has several prices, and only one of them identifies the plan. Given the full item
list from a webhook:
1. Find the item whose price ID matches a `kind: "base"` row. That row gives
`deployment`, `tier` and — by which term key matched — `term`.
2. Sum the quantities of items matching that plan's `kind: "limit"` rows.
3. Collect the feature keys of items matching its `kind: "feature"` rows.
4. Any item matching nothing is a configuration error: fail the event loudly so
it lands on the staff dashboard. Guessing a tier from a price we cannot map is
how a customer ends up with the wrong licence and no record of why.
Only the running `PADDLE_ENV`'s IDs are consulted, so a production process cannot
be talked into resolving a sandbox price. That property is spec 5's and survives
unchanged.
**Out-of-order delivery is still handled by construction.** Paddle sends the
complete item list on every subscription event, so a handler that reads the whole
list is still a function of current state rather than of a transition. Nothing
about metering weakens this.
---
## Issuance
`licensing.Issue` reads the entitlement row for the instance and snapshots
`resolved_limits` and `granted.features`. When no row exists it falls back to the
plan's base — which covers staff manual issuance and any instance predating the
backfill.
`Issue` stays the only signer, and it stays the thing that does not deliver.
**Upgrades preserve the expiry.** A mid-term server addition passes
`ExpiresAt` = the current licence's expiry, so the licence is reissued with a
larger cap and the same end date. It must not extend the term: the customer paid
a prorated amount for the rest of this period, not for a new one. Note that the
current expiry already includes `GracePeriod`, so nothing adds it again —
`ExpiresAt` overriding `Term` is exactly the existing contract.
**Reductions issue nothing.** They live in `desired` with `scheduled_change_at`
set until the renewal webhook promotes `desired` into `granted` and issues the
next term at the lower cap. The customer keeps what they paid for to the end of
the period, there is no refund to reason about, and no licence ever shortens —
which is the rule spec 5 states and this design does not touch.
---
## Changing a live subscription
`PUT /api/instances/:id/entitlement` writes `desired`, then calls Paddle:
- **An increase** updates the subscription items prorated immediately. The
resulting `subscription.updated` webhook promotes `granted` and reissues.
- **A decrease** schedules the item change for the next billing period and sets
`scheduled_change_at`. No licence action now.
This is admin's **first outbound Paddle call beyond the portal session**, and
spec 5 currently states it has none. That statement changes. The important part
does not: **the webhook remains the only thing that promotes `granted` or issues
a licence.** The endpoint writes `desired` and asks Paddle for a change; it never
grants anything itself. A customer whose card is declined on a prorated upgrade
gets no licence, which is correct, and admin needs no compensating logic to
achieve it.
A tier change (Professional to Enterprise) is the same call with a different base
price, and issues with `ReasonTierChange` as it already would.
---
## Control-plane enforcement
**Feature gating already exists and is already mounted.** `RequireFeature` in
`server/internal/api/licence.go` answers 403 `feature_unavailable`, and
`server/internal/api/handlers.go` already wraps `POST /api/console/connect`,
`GET /api/console/tunnel` and `GET`/`PUT /api/org/oidc` in it. Free's feature list
is empty, so a Free instance already cannot open the console. **No capability is
taken away from an existing tenant by this spec, and no customer email is owed.**
**One gap remains, and it is a single check.** `HandleOIDCStart` already tests
`Feature("oidc")` and redirects to `/login?error=oidc_unavailable`.
`HandleOIDCCallback` does not test it at all. A start that 403s is a dead end; an
ungated callback completes a sign-in, so the unguarded half is the half that
matters.
The callback cannot copy the start's instance resolution: the start reads
`InstanceFromHost(c)`, while the callback resolves the instance from the OAuth
state it consumes, and by then it holds `instanceID` directly. The check goes
after `ConsumeStateInstance` and before `providerForInstance`, so a licence that
lapsed mid-flow stops the exchange rather than completing it.
`web/` hides the Console button and the SSO card when the feature is absent, but
as everywhere else in this codebase the API is the boundary and the UI is the
courtesy.
**`CheckMonitorLimit`** joins the three existing checks in
`server/internal/services/licence_limits.go`, counting `monitors` for the
instance. Same shape: refuse a new one at the cap, never truncate what exists.
`LicenseUsage` reports monitors alongside the other counts.
**Audit retention is new work.** Nothing trims `audit_logs` today. A daily sweep
deletes entries older than the licence's `AuditRetentionDays` per instance;
`Unlimited` skips the instance entirely. It is modelled on the existing workflow
log retention sweep, and it is the one item in this design that deletes customer
data — so it must read the *current* licence's value each run rather than caching
it, and an instance whose licence has lapsed must not be swept on the expired
term's allowance.
**Degraded mode is unchanged.** Expiry still stops mutations and leaves monitors
executing, alerts firing and agents keyed. A feature gate is a mutation gate for
console and SSO, so it behaves the same way.
---
## HQ, the configurator
One screen, reached from an instance in `InstanceRecord` and from the
self-hosted purchase page.
```
Deployment ( ) Cloud (•) Self-Hosted ← fixed after creation
Tier ( ) Free (•) Professional ( ) Enterprise
Term (•) Annual ← monthly hidden for self-hosted
Servers [ 10 ] base 3 included, 7 extra
Features [x] Browser console
[ ] Single sign-on
─────────────────────────────────────────────
£B + 7 × £S per year
[ Continue to payment ]
```
It is one component in both places, driven by the catalogue rather than by
anything hardcoded — a feature that gains a price shows its price with no
frontend change, which is the point of the catalogue being data.
**Existing subscriptions show `desired` and `granted` when they differ:** "10
servers, dropping to 5 on 12 August". A pending reduction is a fact about the
account and belongs on the screen, not only in Paddle.
**Choosing Free skips payment entirely.** With no catalogue rows there is no
checkout to open, so the configurator's Continue button links a UUID and issues
directly. For cloud that is the shipped `POST /api/instances`, untouched. For
self-hosted Free it is the existing link flow with no subscription attached — a
new path, and the only place in the system where an instance is licensed without
either a payment or a staff action. It is bounded by the same one-Free-per-account
rule, now scoped per deployment.
**The staff plans editor** edits `plans` (allowances, support level, active) and
`catalogue` (price IDs per environment and term) as two tables. This replaces
spec 5's price-ID editor, which was built for a single map on the plan row.
Follows `adminsite/`'s existing shell without exception: `PageHeader` with its
record line, `PageFrame`'s main-plus-rail split, tokens only and no hex values,
light default. Price and server count read as text as well as position, since
state never reads by colour alone here.
---
## Migration
Admin has no migrations collection: `models.Backfill` runs every boot and is
idempotent by filtering on the absence of what it writes. This all goes there.
1. **Seed six plan rows** from `shared/license/plans.go`, `$setOnInsert` only, so
staff edits to allowances survive a redeploy — the existing `SeedPlans` rule.
2. **Re-key existing plan rows.** The three current rows are keyed by tier alone.
`free` and `professional` gain `deployment: "cloud"`. The row with tier
`self_hosted` becomes `deployment: "self_hosted", tier: "professional"`.
3. **Re-tier existing self-hosted instances and their entitlements.** Instances
holding `tier: "self_hosted"` become `tier: "professional"`; their deployment
already says so.
4. **`license.TierSelfHosted` is kept as a legacy constant** that no new licence
uses. Licences already issued carry `tier: "self_hosted"` in a signed payload
we cannot rewrite, and the server reads limits and features from the payload
rather than from the tier name — so they keep working untouched. This is
exactly what "the server never branches on tier name" was for.
5. **Backfill an entitlement row per instance** from its current licence:
`granted.servers` from `limits.max_servers` (`Unlimited` maps to the plan
base, since an unlimited licence bought no server units), `granted.features`
from the licence's features, `desired` equal to `granted`.
6. **Seed the catalogue** with sixteen rows — the four paid plans times a `base`,
a `limit: max_servers`, a `feature: console` and a `feature: oidc` — price IDs
empty. **The two Free plans get no catalogue rows at all**, which is what keeps
Free outside Paddle: there is nothing to price, so no checkout can be built. Empty price IDs mean checkout refuses until staff paste them, which is
the correct failure: a checkout that silently picks the wrong price is worse
than one that will not open.
Existing licences are not reissued. `MaxMonitors` and `AuditRetentionDays` are
absent from their payloads and decode as `0`, which would read as "no monitors,
trim everything". **Zero must therefore be treated as unset on decode** and
filled from the plan base — a licence signed before a field existed cannot be
allowed to mean the most restrictive possible value of it. This is the one
sharp edge in the whole migration and it is worth a comment at the decode site.
---
## Out of scope
- **Paid feature add-ons.** The model supports one — a `price_ids` entry on a
`kind: "feature"` row — but no feature has a price at launch.
- **Metered channels, monitors or secret groups.** A catalogue row away, and
deliberately not taken.
- **Usage-based billing.** Servers are a configured cap, not a measured count. We
never bill for what an instance ran; we bill for what it is allowed to run.
- **Refunds and credits.** Paddle's, and only Paddle's.
- **Enterprise contract terms, POs and invoicing.** Card only at launch.
- **Anything that revokes or shortens a licence.** Offline verification means
this is not that kind of system, and no part of this design changes it.
---
## Done when
- Six plan rows exist, keyed on `(deployment, tier)`, and a customer can buy any
of the four paid combinations from the configurator.
- A Professional cloud customer can go from 3 to 10 servers and see the new cap
in `web/` without waiting for a renewal.
- The same customer can reduce to 5 and see both the current cap and the date it
drops, with their licence untouched until then.
- Free self-hosted can be created, renewed from HQ, and lapses to read-only
without being reaped.
- Unticking Browser console removes it from the next issued licence, and
`POST /api/console/connect` answers 403 on an instance whose licence lacks it
(already true; the new part is that a customer controls the tick).
- `/auth/oidc/callback` answers 403 on an instance whose licence lacks `oidc`.
- A monitor beyond the cap is refused with a machine-readable 403.
- `audit_logs` older than the licence's retention are gone, and an unlimited
licence's are not.
- Every price ID in the running environment resolves to a plan, and a webhook
naming one that does not fails loudly onto the staff dashboard.
@@ -1,177 +0,0 @@
# Control plane (`web/`) — mobile responsive design
Date: 2026-07-27
Scope: `web/` only. `site/` and `adminsite/` are untouched.
## Problem
`web/` was built for a desktop console and has no mobile handling at all.
- `Sidebar` is a fixed `w-60 h-screen` aside rendered unconditionally by
`app/(app)/layout.tsx`. On a 390px phone it eats 62% of the width.
- Every page opens with `p-8` — 64px of horizontal padding on a screen that has
390px to give.
- Six list pages render 46 column tables. They scroll horizontally, so nothing
overflows the page, but reading a row means swiping.
- The workflow builder is a hard `grid-cols-[1fr_320px]` with `w-[340px]` nodes.
At 390px the inspector alone exceeds the viewport.
- Several grids are unprefixed (`grid-cols-3`, `grid-cols-2`, `grid-cols-4`) and
never collapse.
Next's App Router injects `width=device-width, initial-scale=1` by default, so
the breakpoints *do* fire. This is a layout problem, not a viewport one.
## Decisions
| Decision | Choice | Why |
| --- | --- | --- |
| Sidebar collapse breakpoint | `lg` (< 1024px) | Content is dense — tables plus `lg:grid-cols-3` side rails. Reclaiming 240px helps tablets as much as phones, and `lg` is already where the app's own two-and-three column layouts switch. |
| Table treatment on phones | Card stack below `sm` | Horizontal swiping to read a hostname's status is the single worst thing about the current app on a phone. |
| Workflow builder / console | Best-effort responsive | Usable, not redesigned. No blocking notice — a cramped console beats no console. |
| Verification | Static audit + `next build` + `next lint` | The app is auth-gated behind Mongo, Redis and the Go server; none run in this environment. |
## Design
### 1. The shell
A new client component `web/components/AppShell.tsx` owns the responsive chrome
so `app/(app)/layout.tsx` stays a server component:
```
AppShell (client, holds `open` state)
├── <aside class="hidden lg:flex"> ← permanent sidebar, unchanged look
├── mobile top bar (lg:hidden, sticky, h-14)
│ hamburger · Logo · "Vantage" · instance name
├── offcanvas (lg:hidden, fixed inset-0 z-50)
│ backdrop (bg-black/60) + w-72 panel, translate-x transition
└── <main class="flex-1 overflow-y-auto"> ← LicenseBanner + children
```
`Sidebar.tsx` splits into:
- `SidebarContent` — the nav list, user block and logout. **One copy**, rendered
by both the permanent aside and the offcanvas panel. It takes an optional
`onNavigate` callback so the offcanvas can close on link click.
- `Sidebar` — the permanent `hidden lg:flex` aside.
- `SidebarDrawer` — the offcanvas.
`navItems` and the `activeHref` reduction move to module scope so both
containers share them. The active-item accent bar, the instance name in the
header and the user/logout footer all appear in both, unchanged.
Offcanvas behaviour:
- Closes on route change (`usePathname` effect), on Escape, on backdrop click
and on any nav link click.
- Locks `document.body.style.overflow` while open, restores on close.
- `aria-expanded` / `aria-controls` on the hamburger; `role="dialog"` and
`aria-modal="true"` on the panel; `aria-label` on the button.
- Focus moves into the panel on open and returns to the hamburger on close.
- The panel is always mounted so the slide transition runs in both directions;
it carries `pointer-events-none invisible` when closed rather than being
unmounted.
The top bar is `sticky top-0 z-40` inside the scroll container so it stays
reachable on long pages.
### 2. Tables become card stacks without duplicating markup
The responsive mode lives in the primitives (`web/components/ui/Table.tsx`),
not in each page. Writing two parallel trees per page — a `<table>` for desktop
and a `<div>` stack for mobile — would double six pages of markup and drift
apart on the first edit.
`Td` gains an optional `label`. Below `sm` the table flips to block layout:
| Element | Added classes (below `sm`) |
| --- | --- |
| `Table` | `max-sm:block` |
| `Thead` | `max-sm:hidden` |
| `Tbody` | `max-sm:block max-sm:divide-y-0 max-sm:space-y-3 max-sm:p-3` |
| `Tr` | `max-sm:block max-sm:rounded max-sm:border max-sm:border-border max-sm:bg-surface-2/40 max-sm:p-3` |
| `Td` | `max-sm:flex max-sm:items-start max-sm:justify-between max-sm:gap-4 max-sm:px-0 max-sm:py-1.5` |
When `label` is present, `Td` renders it in a `sm:hidden` span using the exact
mono keyed-label idiom `Th` already uses — `font-mono text-[0.68rem] uppercase
tracking-[0.13em] text-text-secondary`. The key/value pairing on a phone is the
same visual device as the column head on a desktop, because it means the same
thing.
A `Td` with no `label` (the trailing action cell) renders its child alone,
right-aligned in the card.
Pages change only by adding `label="Hostname"` to their cells. Affected:
`servers`, `keys`, `monitors`, `secrets`, `secrets/[group]`, `workflows`,
`workflows/[id]/runs`, `audit`, `keys/[id]`, `servers/[id]` (two tables),
`monitors/[id]`, and `components/settings/MembersCard.tsx`.
### 3. Page padding and headers
- `p-8``p-4 sm:p-6 lg:p-8`, everywhere it opens a page or a page-level
error/loading state — 30 occurrences across 21 files.
- Title-plus-action header rows: `flex items-center justify-between`
`flex flex-col gap-3 sm:flex-row sm:items-center sm:justify-between`. The
action button then sits under the title on a phone rather than squeezing it.
### 4. Modal becomes a bottom sheet under `sm`
`Modal.tsx`: `items-center``items-end sm:items-center`, wrapper `p-4`
`p-0 sm:p-4`, panel gets `rounded-b-none sm:rounded` and `max-h-[85dvh]`
(`dvh`, not `vh` — mobile browser chrome makes `vh` overshoot). Sheets are what
phones expect for a modal, and it costs four classes.
### 5. Workflow builder
Below `lg` the fixed-height two-column grid is dropped entirely: single column,
natural page flow, canvas scrolls with the page.
- `grid h-[calc(100vh-53px)] grid-cols-[1fr_320px]`
`flex flex-col lg:grid lg:h-[calc(100dvh-53px)] lg:grid-cols-[1fr_320px]`.
The viewport-height calculation is `lg:`-only, which matters because the
mobile top bar changes the arithmetic and `100vh` is wrong on mobile anyway.
- Node width `w-[340px]``w-full lg:w-[340px]`; the column wrapper
`w-[340px]``w-full max-w-[340px]`.
- Canvas padding `p-8``p-4 sm:p-6 lg:p-8`.
- The inspector `<aside>` becomes a collapsible bottom panel below `lg`: it
keeps its place in the flex column, gains a top border instead of a left one,
and is hidden until a step is selected (on a phone an empty "Select a step to
configure it" panel is noise).
- The builder's own header row wraps: the action cluster moves to a second line
under `sm`.
### 6. Remaining fixed layouts
| File | Change |
| --- | --- |
| `servers/[id]/page.tsx:164` | `grid-cols-3``grid-cols-2 sm:grid-cols-3` |
| `servers/[id]/page.tsx:495` | install one-liner `min-w-64``min-w-0` so it scrolls internally instead of widening the page |
| `secrets/page.tsx:53` | `grid-cols-2``grid-cols-1 sm:grid-cols-2` |
| `monitors/MonitorForm.tsx:76` | `grid-cols-4``grid-cols-2 sm:grid-cols-4` |
| `monitors/MonitorForm.tsx:98,123,144` | `grid-cols-2``grid-cols-1 sm:grid-cols-2` |
| `workflows/StepPickerModal.tsx:132,168` | `grid-cols-2``grid-cols-1 sm:grid-cols-2` |
| `workflows/[id]/runs/[runId]/page.tsx:254` | matrix table wrapped in `overflow-x-auto`; it is a genuine matrix and stays scrollable |
| `workflows/[id]/runs/[runId]/page.tsx:343` | `p-8``p-4 sm:p-6 lg:p-8` |
| `servers/[id]/console/page.tsx:168` | `p-8``p-4 sm:p-6 lg:p-8`; header/toolbar rows wrap |
The run-detail matrix and the console canvas are the two places that keep
horizontal scrolling. Both are genuinely two-dimensional; stacking them would
destroy the information.
## Non-goals
- No changes to `site/` or `adminsite/`.
- No redesign of the console for touch input (no on-screen keyboard work).
- No new dependencies. Tailwind's `max-sm:` variant and `translate-x` are
enough; no headless-UI or animation library.
- No changes to any API, route or data shape. This is presentation only.
## Verification
1. **Audit** — after the edits, `grep` must return no unprefixed `p-8`,
no unprefixed `grid-cols-[2-9]`, and no `w-[3` fixed node widths outside a
`lg:` prefix in `web/app` and `web/components`.
2. `npx next lint` passes with no new warnings.
3. `npx next build` succeeds.
Screenshot verification is out of scope: the app is auth-gated behind Mongo,
Redis and the Go server, none of which run in this environment.
@@ -1,162 +0,0 @@
# Documentation site — design
**Date:** 2026-07-28
**Status:** approved
## Problem
Vantage has no user-facing documentation. Everything an operator needs — how to
install self-hosted, how to enrol an agent, what a workflow step is, how a
licence gets issued — lives either in `CLAUDE.md` (written for contributors, not
users) or in the code. The marketing site sells the product and the control
plane runs it; neither explains it.
## Solution
A fourth Next-adjacent frontend, `docsite/`, built with Docusaurus v3 and shipped
alongside the marketing site.
### Placement and deployment
- Lives at repo root as `docsite/`.
- Served at **`vantage.hostxtra.co.uk/docs`** — a path on the marketing site's
host, routed by a separate Nginx Proxy Manager custom location rather than by
`site/`. It is a path, not a subdomain, deliberately: `*.vantage.hostxtra.co.uk`
is the per-tenant instance namespace and `APP_ROOT_LABEL` resolves an org from
the label before `vantage`, so a `docs.` label there would be read as a tenant
slug.
This makes `baseUrl: "/docs/"` load-bearing. An NPM custom location forwards
the **full** request path upstream — it does not strip the `/docs` prefix — so
the container serves the build from `/usr/share/nginx/html/docs`, not from the
document root. Prefix, asset URLs and upstream paths then agree with no
rewrite rule to keep in step. Getting this wrong is quiet: the HTML loads and
every stylesheet and script 404s.
- Added to **`deploy/docker-compose.site.yml` only**, published as `3005`.
`deploy/docker-compose.yml` (the self-hosted install) must never mention it,
exactly as it never mentions `site`, `sitesvc`, `admin` or `adminsite`.
- Added to `.gitea/workflows/server-deploy.yml` as a seventh image, rebuilding
on `^docsite/` only, build context `docsite/`.
### Runtime
Docusaurus emits a fully static site, so unlike `site/`, `web/` and
`adminsite/` there is no Node server at runtime. Two stages:
1. `node:26-alpine` builder — `npm ci && npm run build``/app/build`.
2. `nginx:alpine-slim` runner — copies `build/` to
`/usr/share/nginx/html/docs` (see the `/docs` prefix note above), plus a
small `nginx.conf` giving `try_files` a 404 fallback to Docusaurus's
`404.html` and long cache headers on `/docs/assets/`. A bare `/` request
redirects to `/docs/`, so hitting the container directly is not a blank 403.
`nginx:alpine-slim` is roughly 12MB against `caddy:alpine`'s ~50MB, and nothing
here needs Caddy's automatic TLS — the host proxy already terminates it.
Build args, baked at build time the same way `site/`'s are:
| Arg | Purpose |
| --- | --- |
| `DOCS_URL` | site `url`; defaults to `https://vantage.hostxtra.co.uk` |
| `DOCS_BASE_URL` | site `baseUrl`; defaults to `/docs/`. Must match the NPM location and the runner's copy target |
| `APP_URL` | navbar link to the control plane |
| `HQ_URL` | navbar link to the HQ portal |
### Theme
`docsite/src/css/custom.css` carries `site/app/globals.css`'s token blocks
**copied verbatim** — same names, same values — and maps Docusaurus's `--ifm-*`
variables onto them. Docusaurus stamps `data-theme="light|dark"` on `<html>`,
which is the same selector `site/`'s dark block already keys on, so the built-in
toggle works with no extra wiring. Light is the default, matching `site/` and
`adminsite/`.
This makes a **fifth** copy of the token block (`site/`, `adminsite/`, `web/`
dark-only, `shared/mail/templates/layout.html.tmpl` as literal hex, and now
`docsite/`). Nothing enforces the match; `CLAUDE.md`'s Frontend section is
updated to say so. No component in `docsite/` may carry a hex value.
Search is `@easyops-cn/docusaurus-search-local` — index built at compile time,
served from the same origin. No Algolia account, no external host, nothing to
key or rotate.
Docs-only mode: `routeBasePath: "/"`, blog disabled, no tutorial scaffolding.
## Content
Sidebar is authored explicitly in `sidebars.ts` rather than autogenerated, so
ordering is a decision rather than a filename accident.
### Getting Started
| Page | Covers |
| --- | --- |
| `what-is-vantage` | The control plane, the agent, what problem each solves |
| `cloud-vs-self-hosted` | The two deployments, what differs (licensing, HQ-managed users, reaping) |
| `self-hosted-install` | Prereqs (Docker, external MongoDB, DNS, TLS), `docker-compose.yml`, required env, `docker compose up -d` |
| `first-login` | `/setup` bootstrap, first org and owner |
| `first-server` | `POST /servers/new`, the install one-liner, Linux and Windows, watching it flip to `active` |
| `claim-free-licence` | Linking the install to an HQ account, `claim-free` |
`self-hosted-install` is the page the section exists for; it names every
required environment variable with its consequence-of-omission, in particular
`GRPC_HOST` (boot fails, no default is safe) and `KEY_ENCRYPTION_KEY`.
### Vantage (the application)
`servers` (agent install Linux/Windows, inventory, OS updates, agent
self-update) · `ssh-keys` (upload, generate-on-server, assign, revoke, what the
agent writes and when) · `workflows` (step library, default steps and why they
are read-only, the designer, running, live logs, `on_failure`, `output_env`,
workspaces, log retention) · `monitors` (the four check types, server vs agent
runner, retries, incidents, uptime rollups) · `notification-channels` (five
types, testing) · `secrets` (vault, `secret_refs` in steps, the ESO read path
and its bearer token) · `browser-console` (SSH/RDP/VNC, one-time tokens) ·
`audit-log` · `settings` (members and roles, OIDC per org, retention, ESO token,
licence).
### Vantage HQ (the portal)
`accounts-and-signup` (account-first signup, email verification, an account is
a team) · `people-and-roles` (owner/admin/member, invitations, accepting) ·
`cloud-instances` (create, overview, granting members and what a grant actually
is) · `self-hosted-instances` (purchase creates a placeholder, claim-link binds
the real UUID, relink) · `licensing-and-entitlements` (tiers, metered server
count, feature toggles, desired vs granted) · `billing` (Paddle as merchant of
record, checkout, the customer portal, changing configuration) · `free-tier`
(limits, the renewal window, reaping on cloud).
### Reference
`environment-variables` (server, sitesvc, admin, agent) · `rest-api` (the route
tables, grouped as in `CLAUDE.md`) · `grpc-api` (the eight RPCs, the command
stream) · `agent-config` (config.yaml, paths, permissions) ·
`ports-and-networking` (which ports, which direction, what needs to be
reachable) · `troubleshooting`.
### Operations
`upgrading` (pull and recreate) · `backups` (MongoDB is the durable state; Redis
is sessions only) · `agent-updates` · `ci-cd` (which image rebuilds when, and
the repo-variable gap that pushes no commit).
## Writing rules
- Every guide is task-shaped: numbered steps, real paths and commands taken from
the repository, never invented UI.
- Behaviour that is a hard refusal gets an admonition, not a paragraph: default
steps are read-only (409 `ErrDefaultStep`), `POST /license` answers 409
`cloud_managed` on cloud, HQ-sourced users cannot have their role changed
locally (409 `ErrHQManaged`).
- Where the UI enforces something, say that the API is the boundary and the UI
is the courtesy — the same phrasing the codebase uses.
- No screenshots in this pass. They rot faster than prose and there is no
capture pipeline.
## Out of scope
- Versioned documentation. One version, tracking `main`. Docusaurus versioning
can be switched on later without restructuring.
- Internationalisation.
- Screenshots and diagrams beyond what Mermaid renders inline.
- A docs search backed by an external service.
@@ -1,209 +0,0 @@
# Agent-relayed console proxy
Date: 2026-07-29
Status: approved, not yet implemented
## Problem
`consoleTunnel` builds guacamole parameters from `srv.IPAddress` and hands them
to guacd, which then dials the target itself. On a self-hosted deployment the
control plane and the managed servers share a network, so that works. On Vantage
Cloud they do not: guacd runs on the cloud host and the customer's server is on
an RFC1918 address behind their NAT. Every cloud console session to a private
address fails, for SSH, RDP and VNC alike.
Agents already hold an outbound gRPC connection to the control plane. The fix is
to carry the console's TCP bytes over that existing path rather than asking guacd
to route somewhere it cannot reach.
## Decisions
**Self-relay only.** The agent relays to its own host and nowhere else. It is
never told a hostname; the host is hardcoded to `127.0.0.1` on the agent side and
only the port comes from the server. A jump-host mode (reaching agentless devices
through a neighbouring agent) was rejected: it would give an agent the power to
dial arbitrary addresses on the customer's LAN, and the console today can only
target servers that run an agent anyway.
**A dedicated bidirectional RPC, one stream per TCP connection.** Multiplexing
console bytes onto the existing `CommandStream` was rejected — that stream
already carries control commands and workflow stdout, and an RDP framebuffer
would introduce head-of-line blocking against key sync and step output. A
separate stream also gets connection lifetime, flow control and close semantics
for free instead of needing a hand-rolled connection-ID demux.
**Always proxy, both deployments.** Direct dial is deleted rather than kept as a
self-hosted fast path or a fallback. One code path means one tested code path,
and the cloud path is the one no developer can reproduce locally. A
try-direct-then-fall-back design was rejected outright: it puts a timeout in
front of every private-network session and makes "which path did this session
use" unanswerable from the audit log.
The cost is that the console now requires a live agent, where a self-hosted
deployment could previously reach a server whose agent was down. In practice an
offline agent almost always means an offline host, and the failure is now an
immediate, explicit refusal instead of a hang.
## Architecture
Three parties rendezvous on a single `proxy_id`. Neither guacd nor the agent
changes which direction it dials: guacd still makes an outbound TCP connection,
the agent still only connects outbound to the control plane.
```
consoleTunnel (server)
1. proxy.Open(instance, server_id, port) -> proxy_id + ephemeral listener :N
2. push OpenProxyCmd{proxy_id, port} down the existing CommandStream
3. agent dials 127.0.0.1:port locally, then opens ProxyStream and sends
ProxyOpen{server_id, agent_token, proxy_id}
4. guacd dials PROXY_ADVERTISE_HOST:N (the params it was handed in step 1)
5. registry holds both halves -> io.Copy in both directions
6. either side EOFs -> close listener, close stream, drop the registry entry
```
Steps 3 and 4 race, so a registry entry has two slots and starts piping when the
second one arrives. Both waits share a single 10 second deadline; expiry closes
everything and frees the entry.
The agent dials locally *before* opening the stream, so a refused connection
arrives as an explicit `ProxyClose{reason}` rather than as a hang.
`BuildGuacParams` stops reading `srv.IPAddress` and takes the relay host and port
instead. `IPAddress` remains in use for display and for monitors.
## Wire protocol
Additive only; no existing message changes shape.
```protobuf
rpc ProxyStream(stream ProxyClientMsg) returns (stream ProxyServerMsg);
message OpenProxyCmd { // ServerCommand oneof field 8
string proxy_id = 1;
uint32 port = 2;
}
message ProxyClientMsg {
oneof payload {
ProxyOpen open = 1; // first message only
bytes data = 2;
ProxyClose close = 3;
}
}
message ProxyOpen { string server_id = 1; string agent_token = 2; string proxy_id = 3; }
message ProxyServerMsg { oneof payload { bytes data = 1; ProxyClose close = 2; } }
message ProxyClose { string reason = 1; }
```
Two implementation facts about this repo shape the above. The `pb` packages are
**hand-written Go, not protoc output**`vantage.proto` is documentation, and
both `server/internal/grpc/pb` and `agent/internal/grpc/pb` are edited by hand
and kept in sync manually. And the registered codec is JSON, so a `bytes` field
travels as a base64 string: roughly 33% overhead on relayed traffic. That is
accepted rather than fixed here, because introducing a second codec for one RPC
is a larger change than this feature warrants. Relay chunks are 32 KiB.
## Security
**The agent only ever dials `127.0.0.1`.** The port is the only field it takes
from the server; the host is hardcoded agent-side. A compromised control plane
cannot use an agent to reach anything else on the customer's network. This is the
strongest property in the design and the reason self-relay was chosen.
**`proxy_id` is 32 random bytes, single-use and scoped.** On `ProxyOpen` the
server checks three things together: the agent token hash matches that
`server_id`, the `proxy_id` exists in the registry, and the entry's `server_id`
and `instance_id` match the authenticated agent. Any mismatch closes the stream
without revealing which check failed.
**The listener is the exposed surface and is narrowed four ways.** It binds an
ephemeral port; it lives at most 10 seconds unclaimed; it accepts exactly one
connection and closes immediately afterwards; and the accepted connection's
remote address must resolve to a host named in `GUACD_ADDR`. Without that last
check, any other container on the Docker network could claim the session during
the window.
**Agent-offline is refused early.** `consoleConnect` checks
`srv.Status == "active"` and returns 409 `agent_offline`, rather than letting the
browser open a WebSocket that dies on a deadline.
**Audit.** `console.opened` gains the relay port and `proxy_id`. A relay that
expires or is refused writes `console.proxy_failed` with a reason, so a failed
console session stops being invisible.
Credentials are unchanged. Private keys and RDP passwords travel from the server
to guacd inside the guacamole handshake and never reach the agent. The SSH and
RDP sessions are negotiated end-to-end between guacd and the target daemon, so
the agent relays bytes it cannot read.
## Components
New, server:
| Unit | Responsibility |
| --- | --- |
| `server/internal/proxy/registry.go` | `Open`, `AttachAgent`, `AttachTCP`, expiry sweep. Pure state — no net, no gRPC, testable alone |
| `server/internal/proxy/session.go` | One relay: listener, deadline, the `io.Copy` pair, teardown-once |
| `server/internal/grpc/proxystream.go` | The `ProxyStream` handler: authenticate, then hand the stream to the registry. No relay logic of its own |
New, agent:
| Unit | Responsibility |
| --- | --- |
| `agent/internal/proxy/proxy.go` | `Open(ctx, client, proxyID, port)` — dial loopback, open the stream, pump bytes. No build tags; Linux and Windows share it |
Changed:
- `proto/vantage/v1/vantage.proto`, and both generated pb trees
- `server/internal/services/console.go``BuildGuacParams(srv, relayHost, relayPort, …)`
- `server/internal/api/console.go` — offline pre-check in `consoleConnect`; open the relay before the guacd handshake in `consoleTunnel` and close it in `OnDisconnect`
- `agent/internal/sync/sync.go` — handle `OpenProxyCmd`, one goroutine per proxy
- `deploy/docker-compose.yml`, `deploy/docker-compose.site.yml``PROXY_ADVERTISE_HOST=server`
Two new optional environment variables on the server: `PROXY_ADVERTISE_HOST`
(default `server`, the name guacd resolves the control plane by) and
`PROXY_LISTEN_HOST` (default `0.0.0.0`).
Nothing new is opened on the customer's firewall — the relay rides the agent's
existing outbound gRPC connection.
`docsite/docs/reference/ports-and-networking.md` and
`docsite/docs/vantage/browser-console.md` must say so, and must state the new
requirement that the agent be online.
A secondary benefit beyond cloud: a VNC or RDP service bound only to `127.0.0.1`
is now reachable, where a direct dial from guacd never could be.
## Failure modes
| Failure | Behaviour |
| --- | --- |
| Agent offline at connect | 409 `agent_offline` from `consoleConnect`, before any WebSocket is opened |
| Agent never opens the stream | 10s deadline; listener closed; `console.proxy_failed{reason:"agent_timeout"}`; WebSocket closed with a message the UI surfaces |
| Local dial refused (daemon down, wrong port) | `ProxyClose{reason}` relayed up as the same audit event, reason `dial_refused` |
| guacd never dials | Same deadline path, reason `guacd_timeout` |
| Bad token, unknown or foreign `proxy_id` | Stream closed with no detail leaked; `console.proxy_failed{reason:"rejected"}` |
| Agent process dies mid-session | Stream EOF, relay torn down, console shows a disconnect |
| CommandStream reconnects mid-session | No effect on live sessions — the relay is on its own stream. Only a new `OpenProxyCmd` needs the control stream |
Teardown is guarded by `sync.Once` on both sides: both `io.Copy` goroutines
finish, and whichever finishes second must not double-close.
## Testing
Written test-first.
- `server/internal/proxy/registry_test.go` — the two halves pair in either
order; expiry frees the entry; a second claim on a used `proxy_id` is
rejected; a mismatched `instance_id` is rejected. No network.
- `server/internal/proxy/session_test.go` — two `net.Pipe` halves; bytes flow
both ways; EOF in each direction tears down; double-close is safe.
- `server/internal/grpc/proxystream_test.go` — the authentication matrix: valid,
wrong token, unknown `proxy_id`, `proxy_id` belonging to another instance.
- `agent/internal/proxy` — a refused dial emits `ProxyClose`; the happy path
echoes bytes.
- End-to-end in `server`: a fake agent plus a `net.Listen` echo server, asserting
bytes traverse listener → registry → stream → echo and back. This is the test
that would have caught the original bug.
Manual verification, in this order: self-hosted SSH (proves no regression),
cloud SSH to a private-network host, cloud RDP to a Windows agent.
-89
View File
@@ -1,89 +0,0 @@
# Vantage Licensing Programme — Spec Index
Build in this order. Specs 0a5 were designed 2026-07-24; spec 6 on 2026-07-26.
| # | Spec | Plan | Status |
|---|---|---|---|
| 0a | [shared-module](2026-07-24-shared-module-design.md) | [plan](../plans/2026-07-24-shared-module.md) | **shipped** |
| 0b | [instance-rename](2026-07-24-instance-rename-design.md) | [plan](../plans/2026-07-24-instance-rename.md) | **shipped**, migration verified on live |
| 1 | [licensing-core](2026-07-24-licensing-core-design.md) | [plan](../plans/2026-07-24-licensing-core.md) | **shipped** |
| 2 | [instance-licensing](2026-07-24-instance-licensing-design.md) | [plan](../plans/2026-07-24-instance-licensing.md) | **shipped**, no grandfathering — existing cloud instances are read-only until admin backfills |
| 3 | [admin-backend](2026-07-24-admin-backend-design.md) | [plan](../plans/2026-07-24-admin-backend.md) | **shipped**, verified end to end against scratch databases |
| 4 | [admin-site](2026-07-24-admin-site-design.md) | — | ready to start |
| 5 | [paddle-billing](2026-07-24-paddle-billing-design.md) | [plan](../plans/2026-07-27-paddle-billing.md) | **shipped (code)** — client, webhooks, checkout, entitlement update and portal built and compiled against spec-7's catalogue/entitlements; signup-migration dropped (done by 6). Live sandbox catalog + end-to-end pass is the operator's step. Old [2026-07-26 plan](../plans/2026-07-26-paddle-billing.md) superseded. |
| 6 | [cloud-instance-creation](2026-07-26-cloud-instance-creation-design.md) | — | ready to start |
| 7 | [metered-licensing](2026-07-26-metered-licensing-design.md) | [plan](../plans/2026-07-26-metered-licensing.md) | **shipped** — staff can configure and issue any of the six plans; no customer can buy one until 5 lands |
Specs 1 and 2 together give working licensing with licences cut by hand with
`lkctl` — no admin service needed. 4 and 5 can run in parallel once 3 lands.
7 lands before 5. It re-keys `plans` on `(deployment, tier)`, moves every Paddle
price ID out of `plans` into a new `catalogue` collection, and adds the
`entitlements` collection that both a subscription and a licence are derived from
— all of which plan 5 builds on top of, so building 5 first would mean writing
its billing code twice.
## The shape
```
Account (admin only)
├── Instance 1 cloud vantage.hostxtra.co.uk/<slug> licence auto-injected
├── Instance 2 cloud licence auto-injected
└── Instance 3 self-hosted customer's own deployment licence pasted by hand
```
The control plane knows only **Instance**. Accounts exist solely in the admin
service, because a self-hosted instance has no row in the cloud database at all.
## Decisions that everything else follows from
**Licences are offline-verified signed blobs.** ECDSA P-384 with SHA-256 via
`github.com/hyperboloide/lk`, public key compiled into the server, no phone-home
anywhere. This buys air-gapped self-hosting and means no Vantage instance ever
depends on the licensing service being up. It costs revocation: a licence is
valid until it expires whatever Paddle later says. Self Hosted is annual-only to
bound that window.
**Every licence is bound to one instance UUID.** Self-hosted customers link their
UUID before the licence is signed, so there is no unbound licence and no claim
protocol.
**Expiry degrades, it does not break.** Monitors keep executing, alerts keep
firing, agents keep their keys, in-flight workflow runs finish. Mutations stop.
Deletes and OS-update application stay open so a customer is never trapped
over-limit or unpatched.
**Tiers are data, not code.** The server reads `Limits` and `Features` and never
branches on tier name. Tier contents live in the admin `plans` table and are
snapshotted into each issued licence, so editing a plan never rewrites history —
the same rule as `workflow_runs.steps_snapshot`.
Spec 7 replaces the three-tier table below with two deployments times three
tiers, and makes the server count a metered quantity rather than a fixed
allowance. See [metered-licensing](2026-07-26-metered-licensing-design.md) for
the current grid. As shipped through spec 3, the table is:
| | Free | Professional | Self Hosted |
|---|---|---|---|
| deployment | cloud only | cloud | self-hosted |
| max servers | 3 | unlimited | unlimited |
| max secret groups | 1 | unlimited | unlimited |
| max channels | 1 | unlimited | unlimited |
| console | no | yes | yes |
| OIDC | no | yes | yes |
| term | monthly, £0 | monthly or annual | annual only |
Free is cloud-only by construction: it is only ever signed with
`deployment: "cloud"`, and verification rejects a deployment mismatch. There is
no server-side flag to edit. One Free instance per account.
**Spec 7 ends that construction-level guarantee** — there is a self-hosted Free
plan, so `plan.Deployment != inst.Deployment` no longer implies it, and the Free
limit becomes one per account *per deployment*.
**Existing cloud tenants are not grandfathered.** The migration that would have
done it was removed before plan 2 shipped, so every existing cloud instance is
read-only until it is licensed by hand through the admin service: attach it to an
account with `POST /api/staff/instances`, then `POST /api/staff/instances/:id/issue`.
That flow is verified in plan 3, so it works today via the API and is the first
job the admin UI is used for.