Compare commits

...
Author SHA1 Message Date
mrhid6 2bc15648b7 docs: Describe the Windows agent's update and workload support
Chart Release / chart (push) Successful in 34s
Server Deploy / deploy (push) Failing after 5m34s
Agent Release / build (push) Successful in 12m50s
Agent Release / msi (push) Successful in 4m22s
2026-08-13 12:23:34 +00:00
mrhid6 7f348b7b2b fix: Follow Windows Th/Td labels through the updates table 2026-08-13 12:19:47 +00:00
mrhid6 a71fd9a9f4 fix: Fixed docs entitlement 2026-08-13 12:18:21 +00:00
mrhid6 36848a519a feat: Word the workload and update panels for Windows servers 2026-08-13 12:16:10 +00:00
mrhid6 ea6d0b969c fix: Trim Windows event log after SCM filtering, not before 2026-08-13 12:12:11 +00:00
mrhid6 4ba9983cf8 feat: Control Windows services and read their event log as workloads 2026-08-13 12:04:49 +00:00
mrhid6 8caaa4540f fix: Respect .exe word boundary and unterminated quotes in servicePath 2026-08-13 12:00:59 +00:00
mrhid6 42f0cb67cb feat: Collect Windows services as workloads 2026-08-13 11:52:08 +00:00
mrhid6 1dfb3cc28c refactor: Split the agent workloads package by build tag 2026-08-13 10:49:24 +00:00
mrhid6 b51e87477e feat: Report whether a managed host is waiting on a reboot 2026-08-13 10:39:27 +00:00
mrhid6 05c0ad43d2 feat: Check and apply Windows updates through the Windows Update COM API 2026-08-13 10:36:22 +00:00
mrhid6 4a07af049c refactor: Split the agent updates package by build tag 2026-08-13 10:33:35 +00:00
mrhid6 057c193e26 feat: Add winexec helper for running PowerShell from the agent 2026-08-13 10:26:31 +00:00
mrhid6 38a731f269 docs: Add implementation plan for Windows agent parity 2026-08-13 10:21:04 +00:00
mrhid6 06d69590fc docs: Design for Windows agent parity on updates and workloads 2026-08-13 10:08:25 +00:00
mrhid6 6adee810dc feat: Cleanup old docs 2026-08-13 09:41:32 +00:00
53 changed files with 3665 additions and 14983 deletions
+36 -7
View File
@@ -315,6 +315,17 @@ a Service cannot address the one pod holding a console listener.
Agents report CPU/memory/swap/partitions/kernel — metrics every 30s, full static snapshot every 15 min. They also check for pending OS package updates hourly and can apply them on command (`ApplyUpdatesCmd`).
Windows update checking and applying go through the Windows Update COM API
(`Microsoft.Update.Session`) rather than the PSWindowsUpdate module, which would
need a PowerShell Gallery install on every host and fails on an air-gapped
fleet. `CurrentVersion` is empty on Windows and `NewVersion` carries the KB
article ID: a Windows update is not a version bump of a named package.
**The agent never reboots a host.** `ApplyUpdatesCmd` installs and stops there;
`inventory.reboot_required` reports that one is owed, set on the static snapshot
every 15 minutes. Linux fills it too, from `/var/run/reboot-required` or
`dnf needs-restarting -r`.
### Package inventory and CVE findings
Agents report their installed packages hourly; the control plane matches them
@@ -364,10 +375,22 @@ scheduler off entirely.
A **workload** is one Docker container or one systemd unit — one word for the
page, the collection and the commands, rather than saying "container or
service" in every identifier. Linux only, and **not gated by licence**: this
reads as core fleet management, so v1 ships everywhere with no `HasFeature`
check. If that changes the check belongs at `ReportWorkloads`, gating collection
rather than display, exactly as sub-project A does.
service" in every identifier.
On Windows a workload is a Docker container or a Windows **service**, reported
under the same `unit` kind and the same `systemd_ok` / `systemd_error` fields —
one wire shape, worded per platform in the UI, which is the only layer that
knows the host's OS. The platform split lives entirely in the agent, as build
tags (`systemd_linux.go` / `services_windows.go` and the matching `control_`
and `logs_` pairs); the control plane is OS-blind and needed no changes.
Windows collection runs PowerShell through `agent/internal/winexec`, and every
script emits JSON that a build-tag-free parser reads, so the parsers are tested
on Linux — the agent module has no Windows CI.
**Not gated by licence**: this reads as core fleet management, so v1 ships
everywhere with no `HasFeature` check. If that changes the check belongs at
`ReportWorkloads`, gating collection rather than display, exactly as sub-project
A does.
Agents collect on a 60-second ticker and report through `ReportWorkloads` with
the **offer-then-send** handshake the package report already uses. The offer is
@@ -391,8 +414,9 @@ reads get `WorkloadLogsResult`. `CommandStream` republishes **every**
no-op, so this costs nothing and avoids a second result path.
**The protected set is computed agent-side and enforced agent-side.**
`vantage-agent.service`, plus the container ID read from `/proc/self/cgroup`
should the agent ever run in a container. As with the console relay hardcoding
`vantage-agent.service` on Linux, `VantageAgent` on Windows, plus the container
ID read from `/proc/self/cgroup` should the agent ever run in a container. As
with the console relay hardcoding
`127.0.0.1`, the control plane may name a target but the agent decides what it
will do to itself; a server-side denylist alone would be bypassed by the next
dispatch path someone adds, and the failure is unrecoverable from the UI. The
@@ -1144,7 +1168,12 @@ git push origin main # server + web deploy
- **guacd for console** — protocol handling is Guacamole's problem, not ours; we proxy the WebSocket and manage credentials.
- **`org_id` on every document** — isolation enforced at the query layer, not by separate databases.
- **root only** — manages `/root/.ssh/authorized_keys`; no per-user key management.
- **Windows agents are second-class by design** — register, heartbeat, run steps, report inventory; no `authorized_keys` management.
- **Windows agents cover the fleet-management path** — register, heartbeat, run
steps, report inventory, OS updates through the Windows Update COM API, and
workloads (services plus containers, with control and logs). They still do no
`authorized_keys` management, and no package inventory or CVE matching: the
vulnerability feeds this project uses carry no Windows data, so a Windows host
correctly reports `unsupported` rather than a clean bill of health.
- **Both `server` and `web` scale horizontally** — see "Running more than one server replica" below. `web` holds nothing; `server` holds per-agent state that is routed between replicas over Redis rather than duplicated.
- **Deletion lives in the control plane** — admin sends the warnings because it knows the billing address; the control plane performs the delete because it is the only service that knows which collections carry `instance_id`. Mirroring that list into admin would drift, and a drift there deletes the wrong rows.
+17 -31
View File
@@ -1,5 +1,3 @@
package pb
import (
@@ -83,8 +81,6 @@ type UploadKeyResponse struct {
KeyId string `json:"key_id"`
}
type PackageUpdate struct {
Name string `json:"name"`
CurrentVersion string `json:"current_version,omitempty"`
@@ -99,8 +95,6 @@ type ReportUpdatesRequest struct {
type ReportUpdatesResponse struct{}
type CPUReport struct {
Model string `json:"model,omitempty"`
Cores int `json:"cores,omitempty"`
@@ -119,20 +113,19 @@ type PartitionReport struct {
UsedBytes uint64 `json:"used_bytes"`
}
type InventoryReport struct {
ServerId string `json:"server_id"`
AgentToken string `json:"agent_token"`
IncludeStatic bool `json:"include_static"`
CPU *CPUReport `json:"cpu,omitempty"`
Memory *MemReport `json:"memory,omitempty"`
SwapTotal uint64 `json:"swap_total"`
SwapUsed uint64 `json:"swap_used"`
Partitions []PartitionReport `json:"partitions,omitempty"`
Kernel string `json:"kernel,omitempty"`
ServerId string `json:"server_id"`
AgentToken string `json:"agent_token"`
IncludeStatic bool `json:"include_static"`
CPU *CPUReport `json:"cpu,omitempty"`
Memory *MemReport `json:"memory,omitempty"`
SwapTotal uint64 `json:"swap_total"`
SwapUsed uint64 `json:"swap_used"`
Partitions []PartitionReport `json:"partitions,omitempty"`
Kernel string `json:"kernel,omitempty"`
RebootRequired bool `json:"reboot_required,omitempty"`
}
type InventoryReportResponse struct{}
type MonitorSpec struct {
MonitorId string `json:"monitor_id"`
Type string `json:"type"`
@@ -217,8 +210,6 @@ type ServerCommand struct {
// keepalive is not sufficient on its own.
type PingCmd struct{}
type CleanupWorkspaceCmd struct {
WorkspaceId string `json:"workspace_id"`
}
@@ -241,12 +232,12 @@ type GenerateKeyCmd struct {
}
type AgentMessage struct {
ServerId string `json:"server_id"`
AgentToken string `json:"agent_token"`
Ready *AgentReady `json:"ready,omitempty"`
Result *CommandResult `json:"result,omitempty"`
StepResult *StepResult `json:"step_result,omitempty"`
StepOutput *StepOutputChunk `json:"step_output,omitempty"`
ServerId string `json:"server_id"`
AgentToken string `json:"agent_token"`
Ready *AgentReady `json:"ready,omitempty"`
Result *CommandResult `json:"result,omitempty"`
StepResult *StepResult `json:"step_result,omitempty"`
StepOutput *StepOutputChunk `json:"step_output,omitempty"`
WorkloadLogsResult *WorkloadLogsResult `json:"workload_logs_result,omitempty"`
}
@@ -264,8 +255,7 @@ type RunStepCmd struct {
Script string `json:"script"`
Env map[string]string `json:"env,omitempty"`
TimeoutSeconds int `json:"timeout_seconds,omitempty"`
WorkspaceId string `json:"workspace_id,omitempty"`
}
@@ -284,8 +274,6 @@ type StepOutputChunk struct {
Eof bool `json:"eof,omitempty"`
}
type Vantage_CommandStreamClient interface {
Send(*AgentMessage) error
Recv() (*ServerCommand, error)
@@ -308,8 +296,6 @@ func (c *vantageCommandStreamClient) Recv() (*ServerCommand, error) {
return m, nil
}
type Vantage_CommandStreamServer interface {
Send(*ServerCommand) error
Recv() (*AgentMessage, error)
+10
View File
@@ -450,6 +450,16 @@ func runInventory(ctx context.Context, cfg *config.Config) {
r := inventory.Collect(static)
r.ServerId = cfg.ServerID
r.AgentToken = cfg.AgentToken
// Static snapshots only — every 15 minutes, not every 30 seconds. On
// Windows this spawns a PowerShell process, which is not something to
// do twice a minute forever, and a host rebooted by hand clearing the
// flag within a quarter of an hour is soon enough.
//
// Computed here rather than inside inventory.Collect so the inventory
// package gains no dependency on updates.
if static {
r.RebootRequired = updates.RebootRequired()
}
if err := client.ReportInventory(r); err != nil {
log.Printf("report inventory: %v", err)
}
-9
View File
@@ -3,7 +3,6 @@ package agentsync
import (
"context"
"log"
"runtime"
"time"
"gitea.hostxtra.co.uk/mrhid6/vantage/agent/internal/config"
@@ -18,10 +17,6 @@ const workloadInterval = 60 * time.Second
// runWorkloads reports what this host runs, on its own ticker.
func runWorkloads(ctx context.Context, cfg *config.Config) {
if runtime.GOOS != "linux" {
return
}
reportWorkloads(cfg)
ticker := time.NewTicker(workloadInterval)
@@ -42,10 +37,6 @@ func runWorkloads(ctx context.Context, cfg *config.Config) {
// This is the ONLY writer of the server_workloads collection. RefreshWorkloadsCmd
// calls straight into here rather than answering with data of its own.
func reportWorkloads(cfg *config.Config) {
if runtime.GOOS != "linux" {
return
}
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
+14 -228
View File
@@ -1,238 +1,24 @@
package updates
import (
"bufio"
"bytes"
"context"
"os/exec"
"strings"
"time"
)
// PackageUpdate is one pending update. On Linux it is a package with a version
// on each side. On Windows CurrentVersion is empty and NewVersion carries the
// KB article ID: a Windows update is not a version bump of a named package,
// and inventing a current version would put a wrong string in front of an
// operator.
type PackageUpdate struct {
Name string
CurrentVersion string
NewVersion string
}
func detectPM() string {
for _, pm := range []string{"apt-get", "dnf", "yum", "pacman", "zypper", "apk"} {
if _, err := exec.LookPath(pm); err == nil {
if pm == "apt-get" {
return "apt"
}
return pm
}
}
return ""
}
// CheckAvailable lists pending OS updates.
func CheckAvailable() ([]PackageUpdate, error) { return checkAvailable() }
// ApplyAll installs every pending update. It never reboots: a control plane
// silently restarting a production server is unrecoverable from the UI, so the
// reboot stays a decision a person or a workflow makes. RebootRequired reports
// when one is owed.
func ApplyAll() error { return applyAll() }
func CheckAvailable() ([]PackageUpdate, error) {
switch detectPM() {
case "apt":
return checkApt()
case "dnf":
return checkDnfYum("dnf")
case "yum":
return checkDnfYum("yum")
case "pacman":
return checkPacman()
case "zypper":
return checkZypper()
case "apk":
return checkApk()
default:
return nil, nil
}
}
func ApplyAll() error {
switch detectPM() {
case "apt":
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute)
defer cancel()
if err := exec.CommandContext(ctx, "apt-get", "update", "-qq").Run(); err != nil {
return err
}
return exec.CommandContext(ctx, "apt-get", "upgrade", "-y").Run()
case "dnf":
return exec.Command("dnf", "upgrade", "-y").Run()
case "yum":
return exec.Command("yum", "upgrade", "-y").Run()
case "pacman":
return exec.Command("pacman", "-Syu", "--noconfirm").Run()
case "zypper":
return exec.Command("zypper", "update", "-y").Run()
case "apk":
return exec.Command("apk", "upgrade").Run()
default:
return nil
}
}
func checkApt() ([]PackageUpdate, error) {
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
exec.CommandContext(ctx, "apt-get", "update", "-qq").Run()
out, err := exec.Command("apt", "list", "--upgradable").Output()
if err != nil {
return nil, err
}
var updates []PackageUpdate
scanner := bufio.NewScanner(bytes.NewReader(out))
for scanner.Scan() {
line := scanner.Text()
if !strings.Contains(line, "[upgradable from:") {
continue
}
parts := strings.Fields(line)
if len(parts) < 2 {
continue
}
name := strings.SplitN(parts[0], "/", 2)[0]
newVer := parts[1]
oldVer := ""
if idx := strings.Index(line, "upgradable from: "); idx != -1 {
rest := line[idx+len("upgradable from: "):]
oldVer = strings.TrimSuffix(strings.TrimSpace(rest), "]")
}
updates = append(updates, PackageUpdate{Name: name, CurrentVersion: oldVer, NewVersion: newVer})
}
return updates, nil
}
func checkDnfYum(pm string) ([]PackageUpdate, error) {
cmd := exec.Command(pm, "check-update")
out, err := cmd.Output()
if exitErr, ok := err.(*exec.ExitError); ok && exitErr.ExitCode() == 100 {
err = nil
}
if err != nil {
return nil, err
}
var updates []PackageUpdate
pastHeader := false
scanner := bufio.NewScanner(bytes.NewReader(out))
for scanner.Scan() {
line := scanner.Text()
if !pastHeader {
if strings.TrimSpace(line) == "" {
pastHeader = true
}
continue
}
parts := strings.Fields(line)
if len(parts) < 2 {
continue
}
name := strings.SplitN(parts[0], ".", 2)[0]
updates = append(updates, PackageUpdate{Name: name, NewVersion: parts[1]})
}
return updates, nil
}
func checkPacman() ([]PackageUpdate, error) {
out, _ := exec.Command("pacman", "-Qu").Output()
var updates []PackageUpdate
scanner := bufio.NewScanner(bytes.NewReader(out))
for scanner.Scan() {
parts := strings.Fields(scanner.Text())
if len(parts) < 4 {
continue
}
updates = append(updates, PackageUpdate{Name: parts[0], CurrentVersion: parts[1], NewVersion: parts[3]})
}
return updates, nil
}
func checkZypper() ([]PackageUpdate, error) {
out, err := exec.Command("zypper", "list-updates").Output()
if err != nil {
return nil, err
}
var updates []PackageUpdate
scanner := bufio.NewScanner(bytes.NewReader(out))
for scanner.Scan() {
line := scanner.Text()
if !strings.HasPrefix(line, "v |") && !strings.HasPrefix(line, "i |") {
continue
}
parts := strings.Split(line, "|")
if len(parts) < 5 {
continue
}
updates = append(updates, PackageUpdate{
Name: strings.TrimSpace(parts[2]),
CurrentVersion: strings.TrimSpace(parts[3]),
NewVersion: strings.TrimSpace(parts[4]),
})
}
return updates, nil
}
func checkApk() ([]PackageUpdate, error) {
out, err := exec.Command("apk", "list", "--upgradable").Output()
if err != nil {
return nil, err
}
var updates []PackageUpdate
scanner := bufio.NewScanner(bytes.NewReader(out))
for scanner.Scan() {
line := scanner.Text()
if !strings.Contains(line, "[upgradable") {
continue
}
parts := strings.Fields(line)
if len(parts) < 1 {
continue
}
pkgVer := parts[0]
name := apkName(pkgVer)
newVer := apkVersion(pkgVer)
oldVer := ""
if idx := strings.Index(line, "upgradable from:"); idx != -1 {
rest := strings.TrimSpace(line[idx+len("upgradable from:"):])
rest = strings.TrimSuffix(rest, "]")
oldVer = apkVersion(strings.TrimSpace(rest))
}
updates = append(updates, PackageUpdate{Name: name, CurrentVersion: oldVer, NewVersion: newVer})
}
return updates, nil
}
func apkName(pkgVer string) string {
parts := strings.Split(pkgVer, "-")
var name []string
for _, p := range parts {
if len(p) > 0 && p[0] >= '0' && p[0] <= '9' {
break
}
name = append(name, p)
}
return strings.Join(name, "-")
}
func apkVersion(pkgVer string) string {
parts := strings.Split(pkgVer, "-")
var ver []string
inVer := false
for _, p := range parts {
if !inVer && len(p) > 0 && p[0] >= '0' && p[0] <= '9' {
inVer = true
}
if inVer {
ver = append(ver, p)
}
}
return strings.Join(ver, "-")
}
// RebootRequired reports whether this host is waiting on a restart.
func RebootRequired() bool { return rebootRequired() }
+252
View File
@@ -0,0 +1,252 @@
package updates
import (
"bufio"
"bytes"
"context"
"os"
"os/exec"
"strings"
"time"
)
func detectPM() string {
for _, pm := range []string{"apt-get", "dnf", "yum", "pacman", "zypper", "apk"} {
if _, err := exec.LookPath(pm); err == nil {
if pm == "apt-get" {
return "apt"
}
return pm
}
}
return ""
}
func checkAvailable() ([]PackageUpdate, error) {
switch detectPM() {
case "apt":
return checkApt()
case "dnf":
return checkDnfYum("dnf")
case "yum":
return checkDnfYum("yum")
case "pacman":
return checkPacman()
case "zypper":
return checkZypper()
case "apk":
return checkApk()
default:
return nil, nil
}
}
func applyAll() error {
switch detectPM() {
case "apt":
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute)
defer cancel()
if err := exec.CommandContext(ctx, "apt-get", "update", "-qq").Run(); err != nil {
return err
}
return exec.CommandContext(ctx, "apt-get", "upgrade", "-y").Run()
case "dnf":
return exec.Command("dnf", "upgrade", "-y").Run()
case "yum":
return exec.Command("yum", "upgrade", "-y").Run()
case "pacman":
return exec.Command("pacman", "-Syu", "--noconfirm").Run()
case "zypper":
return exec.Command("zypper", "update", "-y").Run()
case "apk":
return exec.Command("apk", "upgrade").Run()
default:
return nil
}
}
// rebootRequired reads what the distributions themselves record. Debian and
// Ubuntu drop a file; the RPM family answers through needs-restarting, whose
// exit code is 1 when a reboot is owed and 0 when it is not.
func rebootRequired() bool {
if _, err := os.Stat("/var/run/reboot-required"); err == nil {
return true
}
if _, err := exec.LookPath("dnf"); err == nil {
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if err := exec.CommandContext(ctx, "dnf", "needs-restarting", "-r").Run(); err != nil {
if ee, ok := err.(*exec.ExitError); ok && ee.ExitCode() == 1 {
return true
}
}
}
return false
}
func checkApt() ([]PackageUpdate, error) {
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
exec.CommandContext(ctx, "apt-get", "update", "-qq").Run()
out, err := exec.Command("apt", "list", "--upgradable").Output()
if err != nil {
return nil, err
}
var updates []PackageUpdate
scanner := bufio.NewScanner(bytes.NewReader(out))
for scanner.Scan() {
line := scanner.Text()
if !strings.Contains(line, "[upgradable from:") {
continue
}
parts := strings.Fields(line)
if len(parts) < 2 {
continue
}
name := strings.SplitN(parts[0], "/", 2)[0]
newVer := parts[1]
oldVer := ""
if idx := strings.Index(line, "upgradable from: "); idx != -1 {
rest := line[idx+len("upgradable from: "):]
oldVer = strings.TrimSuffix(strings.TrimSpace(rest), "]")
}
updates = append(updates, PackageUpdate{Name: name, CurrentVersion: oldVer, NewVersion: newVer})
}
return updates, nil
}
func checkDnfYum(pm string) ([]PackageUpdate, error) {
cmd := exec.Command(pm, "check-update")
out, err := cmd.Output()
if exitErr, ok := err.(*exec.ExitError); ok && exitErr.ExitCode() == 100 {
err = nil
}
if err != nil {
return nil, err
}
var updates []PackageUpdate
pastHeader := false
scanner := bufio.NewScanner(bytes.NewReader(out))
for scanner.Scan() {
line := scanner.Text()
if !pastHeader {
if strings.TrimSpace(line) == "" {
pastHeader = true
}
continue
}
parts := strings.Fields(line)
if len(parts) < 2 {
continue
}
name := strings.SplitN(parts[0], ".", 2)[0]
updates = append(updates, PackageUpdate{Name: name, NewVersion: parts[1]})
}
return updates, nil
}
func checkPacman() ([]PackageUpdate, error) {
out, _ := exec.Command("pacman", "-Qu").Output()
var updates []PackageUpdate
scanner := bufio.NewScanner(bytes.NewReader(out))
for scanner.Scan() {
parts := strings.Fields(scanner.Text())
if len(parts) < 4 {
continue
}
updates = append(updates, PackageUpdate{Name: parts[0], CurrentVersion: parts[1], NewVersion: parts[3]})
}
return updates, nil
}
func checkZypper() ([]PackageUpdate, error) {
out, err := exec.Command("zypper", "list-updates").Output()
if err != nil {
return nil, err
}
var updates []PackageUpdate
scanner := bufio.NewScanner(bytes.NewReader(out))
for scanner.Scan() {
line := scanner.Text()
if !strings.HasPrefix(line, "v |") && !strings.HasPrefix(line, "i |") {
continue
}
parts := strings.Split(line, "|")
if len(parts) < 5 {
continue
}
updates = append(updates, PackageUpdate{
Name: strings.TrimSpace(parts[2]),
CurrentVersion: strings.TrimSpace(parts[3]),
NewVersion: strings.TrimSpace(parts[4]),
})
}
return updates, nil
}
func checkApk() ([]PackageUpdate, error) {
out, err := exec.Command("apk", "list", "--upgradable").Output()
if err != nil {
return nil, err
}
var updates []PackageUpdate
scanner := bufio.NewScanner(bytes.NewReader(out))
for scanner.Scan() {
line := scanner.Text()
if !strings.Contains(line, "[upgradable") {
continue
}
parts := strings.Fields(line)
if len(parts) < 1 {
continue
}
pkgVer := parts[0]
name := apkName(pkgVer)
newVer := apkVersion(pkgVer)
oldVer := ""
if idx := strings.Index(line, "upgradable from:"); idx != -1 {
rest := strings.TrimSpace(line[idx+len("upgradable from:"):])
rest = strings.TrimSuffix(rest, "]")
oldVer = apkVersion(strings.TrimSpace(rest))
}
updates = append(updates, PackageUpdate{Name: name, CurrentVersion: oldVer, NewVersion: newVer})
}
return updates, nil
}
func apkName(pkgVer string) string {
parts := strings.Split(pkgVer, "-")
var name []string
for _, p := range parts {
if len(p) > 0 && p[0] >= '0' && p[0] <= '9' {
break
}
name = append(name, p)
}
return strings.Join(name, "-")
}
func apkVersion(pkgVer string) string {
parts := strings.Split(pkgVer, "-")
var ver []string
inVer := false
for _, p := range parts {
if !inVer && len(p) > 0 && p[0] >= '0' && p[0] <= '9' {
inVer = true
}
if inVer {
ver = append(ver, p)
}
}
return strings.Join(ver, "-")
}
+9
View File
@@ -0,0 +1,9 @@
//go:build !linux && !windows
// The build constraint above is load-bearing: "_other" is not a GOOS suffix, so
// without it this file compiles on Linux too and collides with updates_linux.go.
package updates
func checkAvailable() ([]PackageUpdate, error) { return nil, nil }
func applyAll() error { return nil }
func rebootRequired() bool { return false }
+117
View File
@@ -0,0 +1,117 @@
package updates
import (
"context"
"fmt"
"strings"
"time"
"gitea.hostxtra.co.uk/mrhid6/vantage/agent/internal/winexec"
)
const (
// The first search after a boot contacts Microsoft Update (or WSUS) and is
// routinely slow. Ten minutes is not generous, it is realistic.
searchTimeout = 10 * time.Minute
// A patch-Tuesday cumulative genuinely takes this long to download and
// install on a modest server.
applyTimeout = 60 * time.Minute
rebootTimeout = 2 * time.Minute
)
// The Windows Update COM API is used rather than the PSWindowsUpdate module: it
// is present on every supported Windows, needs no PowerShell Gallery install,
// and works unchanged against a WSUS server on an air-gapped fleet. The agent
// runs as LocalSystem, which holds the rights it requires.
const searchScript = `
$ErrorActionPreference = 'Stop'
$searcher = (New-Object -ComObject Microsoft.Update.Session).CreateUpdateSearcher()
$result = $searcher.Search("IsInstalled=0 and Type='Software' and IsHidden=0")
$rows = @()
foreach ($u in $result.Updates) {
$ids = @($u.KBArticleIDs)
$kb = ''
if ($ids.Count -gt 0) { $kb = [string]$ids[0] }
$rows += [pscustomobject]@{ title = [string]$u.Title; kb = $kb }
}
ConvertTo-Json -InputObject @($rows) -Depth 3 -Compress
`
const applyScript = `
$ErrorActionPreference = 'Stop'
$session = New-Object -ComObject Microsoft.Update.Session
$result = $session.CreateUpdateSearcher().Search("IsInstalled=0 and Type='Software' and IsHidden=0")
$batch = New-Object -ComObject Microsoft.Update.UpdateColl
foreach ($u in $result.Updates) {
if ($u.InstallationBehavior.CanRequestUserInput) { continue }
if (-not $u.EulaAccepted) {
try { $u.AcceptEula() } catch { continue }
}
$null = $batch.Add($u)
}
if ($batch.Count -eq 0) { Write-Output 'nothing-to-install'; exit 0 }
$downloader = $session.CreateUpdateDownloader()
$downloader.Updates = $batch
$null = $downloader.Download()
$installer = $session.CreateUpdateInstaller()
$installer.Updates = $batch
$r = $installer.Install()
Write-Output ('resultcode=' + $r.ResultCode)
# 2 = succeeded, 3 = succeeded with errors. Anything else failed, and this
# process must exit non-zero so the agent logs a failure rather than an ack.
if ($r.ResultCode -ne 2 -and $r.ResultCode -ne 3) { exit 1 }
exit 0
`
const rebootScript = `
$ErrorActionPreference = 'SilentlyContinue'
$si = New-Object -ComObject Microsoft.Update.SystemInfo
if ($si.RebootRequired) { Write-Output 'true'; exit 0 }
$keys = @(
'HKLM:\SOFTWARE\Microsoft\Windows\CurrentVersion\Component Based Servicing\RebootPending',
'HKLM:\SOFTWARE\Microsoft\Windows\CurrentVersion\WindowsUpdate\Auto Update\RebootRequired'
)
foreach ($k in $keys) { if (Test-Path $k) { Write-Output 'true'; exit 0 } }
$sm = Get-ItemProperty 'HKLM:\SYSTEM\CurrentControlSet\Control\Session Manager' -Name PendingFileRenameOperations
if ($sm -and $sm.PendingFileRenameOperations) { Write-Output 'true'; exit 0 }
Write-Output 'false'
`
func checkAvailable() ([]PackageUpdate, error) {
ctx, cancel := context.WithTimeout(context.Background(), searchTimeout)
defer cancel()
out, err := winexec.Run(ctx, searchScript)
if err != nil {
return nil, fmt.Errorf("windows update search: %w", err)
}
return parseUpdateSearch(out)
}
func applyAll() error {
ctx, cancel := context.WithTimeout(context.Background(), applyTimeout)
defer cancel()
if _, err := winexec.Run(ctx, applyScript); err != nil {
return fmt.Errorf("windows update install: %w", err)
}
return nil
}
func rebootRequired() bool {
ctx, cancel := context.WithTimeout(context.Background(), rebootTimeout)
defer cancel()
out, err := winexec.Run(ctx, rebootScript)
if err != nil {
return false
}
return strings.TrimSpace(out) == "true"
}
+48
View File
@@ -0,0 +1,48 @@
package updates
import (
"encoding/json"
"strings"
)
// winUpdate is one row of the Windows Update searcher's output, in the shape
// searchScript emits it.
type winUpdate struct {
Title string `json:"title"`
KB string `json:"kb"`
}
// parseUpdateSearch reads the searcher's JSON.
//
// It carries no build tag on purpose: this is the half of the Windows update
// path that can be tested on a development machine, and the agent module has no
// Windows CI.
func parseUpdateSearch(jsonText string) ([]PackageUpdate, error) {
s := strings.TrimSpace(jsonText)
if s == "" || s == "null" {
return nil, nil
}
var rows []winUpdate
if err := json.Unmarshal([]byte(s), &rows); err != nil {
// ConvertTo-Json renders a one-element array as a bare object.
var one winUpdate
if err2 := json.Unmarshal([]byte(s), &one); err2 != nil {
return nil, err
}
rows = []winUpdate{one}
}
out := make([]PackageUpdate, 0, len(rows))
for _, r := range rows {
u := PackageUpdate{Name: r.Title}
if kb := strings.TrimSpace(r.KB); kb != "" {
if !strings.HasPrefix(strings.ToUpper(kb), "KB") {
kb = "KB" + kb
}
u.NewVersion = kb
}
out = append(out, u)
}
return out, nil
}
+69
View File
@@ -0,0 +1,69 @@
package updates
import "testing"
func TestParseUpdateSearchArray(t *testing.T) {
in := `[{"title":"2026-08 Cumulative Update for Windows Server 2022","kb":"5034123"},
{"title":"Windows Malicious Software Removal Tool","kb":"890830"}]`
got, err := parseUpdateSearch(in)
if err != nil {
t.Fatalf("parseUpdateSearch: %v", err)
}
if len(got) != 2 {
t.Fatalf("got %d updates, want 2", len(got))
}
if got[0].Name != "2026-08 Cumulative Update for Windows Server 2022" {
t.Errorf("Name = %q", got[0].Name)
}
if got[0].NewVersion != "KB5034123" {
t.Errorf("NewVersion = %q, want KB5034123", got[0].NewVersion)
}
if got[0].CurrentVersion != "" {
t.Errorf("CurrentVersion = %q, want empty", got[0].CurrentVersion)
}
}
// PowerShell 5.1's ConvertTo-Json collapses a one-element array into a bare
// object. A host with exactly one pending update is common, and a parser that
// only accepts arrays reports it as zero.
func TestParseUpdateSearchSingleObject(t *testing.T) {
got, err := parseUpdateSearch(`{"title":"Security Intelligence Update","kb":"2267602"}`)
if err != nil {
t.Fatalf("parseUpdateSearch: %v", err)
}
if len(got) != 1 || got[0].NewVersion != "KB2267602" {
t.Fatalf("got %+v", got)
}
}
func TestParseUpdateSearchNoKB(t *testing.T) {
got, err := parseUpdateSearch(`[{"title":"Driver update for Contoso NIC","kb":""}]`)
if err != nil {
t.Fatalf("parseUpdateSearch: %v", err)
}
if len(got) != 1 || got[0].NewVersion != "" {
t.Fatalf("got %+v, want one update with an empty NewVersion", got)
}
}
// An empty result set is "nothing pending", not a parse failure.
func TestParseUpdateSearchEmpty(t *testing.T) {
for _, in := range []string{"", " \r\n", "[]", "null"} {
got, err := parseUpdateSearch(in)
if err != nil {
t.Fatalf("parseUpdateSearch(%q): %v", in, err)
}
if len(got) != 0 {
t.Fatalf("parseUpdateSearch(%q) = %+v, want none", in, got)
}
}
}
// A KB already carrying its prefix must not become KBKB5034123.
func TestParseUpdateSearchPrefixedKB(t *testing.T) {
got, _ := parseUpdateSearch(`[{"title":"x","kb":"KB5034123"}]`)
if got[0].NewVersion != "KB5034123" {
t.Fatalf("NewVersion = %q", got[0].NewVersion)
}
}
+24
View File
@@ -0,0 +1,24 @@
// Package winexec runs PowerShell on Windows hosts.
//
// It exists because three subsystems — updates, workload collection and
// workload logs — all need the same invocation, and because getting a
// multi-line script past Go quoting, cmd.exe quoting and PowerShell's own
// parser is a problem worth solving once.
package winexec
import (
"encoding/base64"
"unicode/utf16"
)
// EncodeCommand renders a script for powershell.exe -EncodedCommand: UTF-16LE,
// no byte-order mark, base64. This is deliberately free of build tags so it is
// tested on a Linux development machine like every other pure function here.
func EncodeCommand(script string) string {
units := utf16.Encode([]rune(script))
b := make([]byte, 0, len(units)*2)
for _, u := range units {
b = append(b, byte(u), byte(u>>8))
}
return base64.StdEncoding.EncodeToString(b)
}
+20
View File
@@ -0,0 +1,20 @@
package winexec
import "testing"
func TestEncodeCommand(t *testing.T) {
// "hi" as UTF-16LE is 68 00 69 00, which base64-encodes to aABpAA==.
if got := EncodeCommand("hi"); got != "aABpAA==" {
t.Fatalf("EncodeCommand(hi) = %q, want aABpAA==", got)
}
}
func TestEncodeCommandMultiline(t *testing.T) {
// Only that it round-trips through the same encoding PowerShell expects:
// every ASCII byte followed by a zero byte, no BOM.
got := EncodeCommand("a\nb")
want := "YQAKAGIA"
if got != want {
t.Fatalf("EncodeCommand = %q, want %q", got, want)
}
}
+31
View File
@@ -0,0 +1,31 @@
package winexec
import (
"context"
"fmt"
"os/exec"
"strings"
)
// Run executes a PowerShell script and returns its stdout.
//
// powershell.exe rather than pwsh: everything this agent runs through here
// touches Windows Update COM or CIM, both of which are most reliable under
// Windows PowerShell 5.1, and 5.1 is present on every supported Windows while
// pwsh is an optional install.
func Run(ctx context.Context, script string) (string, error) {
cmd := exec.CommandContext(ctx, "powershell.exe",
"-NoProfile", "-NonInteractive", "-EncodedCommand", EncodeCommand(script))
out, err := cmd.Output()
if err != nil {
if ee, ok := err.(*exec.ExitError); ok && len(ee.Stderr) > 0 {
return "", fmt.Errorf("powershell: %s", strings.TrimSpace(string(ee.Stderr)))
}
if ctx.Err() == context.DeadlineExceeded {
return "", fmt.Errorf("powershell: timed out")
}
return "", fmt.Errorf("powershell: %w", err)
}
return string(out), nil
}
+4 -45
View File
@@ -4,9 +4,6 @@ import (
"context"
"errors"
"fmt"
"os"
"os/exec"
"regexp"
"strings"
"time"
)
@@ -14,34 +11,12 @@ import (
// ErrProtected is returned for a workload the agent will not act on.
var ErrProtected = errors.New("workload is protected")
// AgentUnit is the systemd unit this agent runs as.
const AgentUnit = "vantage-agent.service"
// controlTimeout bounds a stop that may never finish on its own. `docker stop`
// waits on a container that may ignore SIGTERM, and `systemctl stop` on a unit
// with a long TimeoutStopSec blocks for exactly as long as that says. A
// waits on a container that may ignore SIGTERM, and both systemctl and
// Stop-Service block for as long as the unit's own stop timeout says. A
// timeout must return a real error rather than an ack implying success.
const controlTimeout = 90 * time.Second
// ownContainerID is read once: the container this agent runs in, if any.
var ownContainerID = detectOwnContainer()
var cgroupContainerRe = regexp.MustCompile(`[0-9a-f]{64}`)
// detectOwnContainer returns this process's container ID, or "" on a host
// install. The agent is normally a systemd service, so "" is the common case;
// this exists so containerising it later cannot silently remove the guard.
func detectOwnContainer() string {
b, err := os.ReadFile("/proc/self/cgroup")
if err != nil {
return ""
}
if m := cgroupContainerRe.FindString(string(b)); m != "" {
return m
}
return ""
}
// isProtected reports whether the agent refuses to act on this workload.
//
// The refusal lives here, in the agent, and not in the control plane. As with
@@ -50,7 +25,7 @@ func detectOwnContainer() string {
// denylist alone would be bypassed by the next dispatch path someone adds.
func isProtected(kind, id, name string) bool {
if kind == "unit" {
return id == AgentUnit || name == strings.TrimSuffix(AgentUnit, ".service")
return isProtectedUnit(id, name)
}
if ownContainerID == "" {
return false
@@ -85,21 +60,5 @@ func Control(ctx context.Context, kind, id, action string) error {
ctx, cancel := context.WithTimeout(ctx, controlTimeout)
defer cancel()
var cmd *exec.Cmd
switch kind {
case "container":
cmd = exec.CommandContext(ctx, "docker", action, id)
case "unit":
cmd = exec.CommandContext(ctx, "systemctl", action, id)
default:
return fmt.Errorf("unknown workload kind %q", kind)
}
if out, err := cmd.CombinedOutput(); err != nil {
if ctx.Err() == context.DeadlineExceeded {
return fmt.Errorf("%s %s timed out after %s", action, id, controlTimeout)
}
return fmt.Errorf("%s %s: %s", action, id, strings.TrimSpace(string(out)))
}
return nil
return controlPlatform(ctx, kind, id, action)
}
+56
View File
@@ -0,0 +1,56 @@
package workloads
import (
"context"
"fmt"
"os"
"os/exec"
"regexp"
"strings"
)
// AgentUnit is the systemd unit this agent runs as.
const AgentUnit = "vantage-agent.service"
// ownContainerID is read once: the container this agent runs in, if any.
var ownContainerID = detectOwnContainer()
var cgroupContainerRe = regexp.MustCompile(`[0-9a-f]{64}`)
// detectOwnContainer returns this process's container ID, or "" on a host
// install. The agent is normally a systemd service, so "" is the common case;
// this exists so containerising it later cannot silently remove the guard.
func detectOwnContainer() string {
b, err := os.ReadFile("/proc/self/cgroup")
if err != nil {
return ""
}
if m := cgroupContainerRe.FindString(string(b)); m != "" {
return m
}
return ""
}
func isProtectedUnit(id, name string) bool {
return id == AgentUnit || name == strings.TrimSuffix(AgentUnit, ".service")
}
func controlPlatform(ctx context.Context, kind, id, action string) error {
var cmd *exec.Cmd
switch kind {
case "container":
cmd = exec.CommandContext(ctx, "docker", action, id)
case "unit":
cmd = exec.CommandContext(ctx, "systemctl", action, id)
default:
return fmt.Errorf("unknown workload kind %q", kind)
}
if out, err := cmd.CombinedOutput(); err != nil {
if ctx.Err() == context.DeadlineExceeded {
return fmt.Errorf("%s %s timed out after %s", action, id, controlTimeout)
}
return fmt.Errorf("%s %s: %s", action, id, strings.TrimSpace(string(out)))
}
return nil
}
@@ -0,0 +1,74 @@
package workloads
import (
"context"
"fmt"
"os/exec"
"strings"
"gitea.hostxtra.co.uk/mrhid6/vantage/agent/internal/winexec"
)
// AgentUnit is the service this agent runs as — the NSSM service name written
// by installer/setup.ps1. Change one, change the other.
const AgentUnit = "VantageAgent"
// A Windows agent is never itself in a container; the Linux build reads
// /proc/self/cgroup, and there is no equivalent question to ask here.
var ownContainerID = ""
// Windows service names are case-insensitive, so the comparison must be too.
func isProtectedUnit(id, name string) bool {
return strings.EqualFold(id, AgentUnit) || strings.EqualFold(name, AgentUnit)
}
func controlPlatform(ctx context.Context, kind, id, action string) error {
switch kind {
case "container":
// Docker behaves identically on Windows, so this path is shared in
// spirit with the Linux one rather than routed through PowerShell.
cmd := exec.CommandContext(ctx, "docker", action, id)
if out, err := cmd.CombinedOutput(); err != nil {
if ctx.Err() == context.DeadlineExceeded {
return fmt.Errorf("%s %s timed out after %s", action, id, controlTimeout)
}
return fmt.Errorf("%s %s: %s", action, id, strings.TrimSpace(string(out)))
}
return nil
case "unit":
// -Force is required: Stop-Service without it refuses outright when
// another service depends on the target, and that refusal reads to an
// operator as a silent no-op.
//
// sc.exe is avoided because it returns before the operation completes,
// which turns a timeout into a false success.
var verb string
switch action {
case "start":
verb = "Start-Service"
case "stop":
verb = "Stop-Service"
case "restart":
verb = "Restart-Service"
default:
return fmt.Errorf("unknown action %q", action)
}
script := "$ErrorActionPreference='Stop'\n" + verb + " -Name " + psQuote(id)
if action != "start" {
script += " -Force"
}
if _, err := winexec.Run(ctx, script); err != nil {
if ctx.Err() == context.DeadlineExceeded {
return fmt.Errorf("%s %s timed out after %s", action, id, controlTimeout)
}
return fmt.Errorf("%s %s: %w", action, id, err)
}
return nil
default:
return fmt.Errorf("unknown workload kind %q", kind)
}
}
+4 -21
View File
@@ -2,9 +2,6 @@ package workloads
import (
"context"
"fmt"
"os/exec"
"strconv"
"strings"
"time"
)
@@ -35,26 +32,12 @@ func Logs(ctx context.Context, kind, id string, tail int) (string, bool, error)
ctx, cancel := context.WithTimeout(ctx, logTimeout)
defer cancel()
var cmd *exec.Cmd
switch kind {
case "container":
cmd = exec.CommandContext(ctx, "docker", "logs",
"--tail", strconv.Itoa(tail), "--timestamps", id)
case "unit":
cmd = exec.CommandContext(ctx, "journalctl", "-u", id,
"-n", strconv.Itoa(tail), "--no-pager", "--output=short-iso")
default:
return "", false, fmt.Errorf("unknown workload kind %q", kind)
out, err := logsPlatform(ctx, kind, id, tail)
if err != nil {
return "", false, err
}
// docker logs writes container stderr to our stderr, so both streams must
// be captured or half the output silently disappears.
out, err := cmd.CombinedOutput()
if err != nil && len(out) == 0 {
return "", false, fmt.Errorf("read logs for %s: %s", id, errText(err))
}
text, truncated := capLog(string(out))
text, truncated := capLog(out)
return text, truncated, nil
}
+30
View File
@@ -0,0 +1,30 @@
package workloads
import (
"context"
"fmt"
"os/exec"
"strconv"
)
func logsPlatform(ctx context.Context, kind, id string, tail int) (string, error) {
var cmd *exec.Cmd
switch kind {
case "container":
cmd = exec.CommandContext(ctx, "docker", "logs",
"--tail", strconv.Itoa(tail), "--timestamps", id)
case "unit":
cmd = exec.CommandContext(ctx, "journalctl", "-u", id,
"-n", strconv.Itoa(tail), "--no-pager", "--output=short-iso")
default:
return "", fmt.Errorf("unknown workload kind %q", kind)
}
// docker logs writes container stderr to our stderr, so both streams must
// be captured or half the output silently disappears.
out, err := cmd.CombinedOutput()
if err != nil && len(out) == 0 {
return "", fmt.Errorf("read logs for %s: %s", id, errText(err))
}
return string(out), nil
}
+89
View File
@@ -0,0 +1,89 @@
package workloads
import (
"context"
"fmt"
"os/exec"
"strconv"
"gitea.hostxtra.co.uk/mrhid6/vantage/agent/internal/winexec"
)
func logsPlatform(ctx context.Context, kind, id string, tail int) (string, error) {
switch kind {
case "container":
cmd := exec.CommandContext(ctx, "docker", "logs",
"--tail", strconv.Itoa(tail), "--timestamps", id)
out, err := cmd.CombinedOutput()
if err != nil && len(out) == 0 {
return "", fmt.Errorf("read logs for %s: %s", id, errText(err))
}
return string(out), nil
case "unit":
display := serviceDisplayName(ctx, id)
// Timestamps are formatted PowerShell-side rather than left to
// ConvertTo-Json, whose DateTime rendering differs between PowerShell
// versions — one of them emits /Date(1699...)/.
//
// $ErrorActionPreference = 'SilentlyContinue' because Get-WinEvent
// treats "no events matched" as a terminating error, and a quiet
// service is normal.
names := psQuote(id)
if display != "" && display != id {
names += "," + psQuote(display)
}
names += "," + psQuote(scmProvider)
// ProviderName includes the host-wide Service Control Manager, so a
// -MaxEvents cap of exactly tail would apply to the combined stream
// before parseEvents narrows SCM rows down to this service — on a
// host with busy service churn the target's own events could be
// squeezed out of the window entirely. Over-fetch instead, hard-capped
// so a pathological host cannot pull an unbounded batch across the
// wire, and let parseEvents trim to the last tail lines after
// filtering.
fetch := tail * 5
if fetch > 2500 {
fetch = 2500
}
script := `
$ErrorActionPreference = 'SilentlyContinue'
$rows = Get-WinEvent -FilterHashtable @{LogName='System','Application'; ProviderName=@(` + names + `)} ` +
`-MaxEvents ` + strconv.Itoa(fetch) + ` |
ForEach-Object {
[pscustomobject]@{
t = $_.TimeCreated.ToUniversalTime().ToString('o')
l = [string]$_.LevelDisplayName
p = [string]$_.ProviderName
m = [string]$_.Message
}
}
ConvertTo-Json -InputObject @($rows) -Depth 3 -Compress
`
out, err := winexec.Run(ctx, script)
if err != nil {
return "", fmt.Errorf("read events for %s: %w", id, err)
}
return parseEvents(out, id, display, tail)
default:
return "", fmt.Errorf("unknown workload kind %q", kind)
}
}
// serviceDisplayName resolves a service's display name, which is what Service
// Control Manager events name it by. An empty answer is fine — the filter then
// matches on the service name alone.
func serviceDisplayName(ctx context.Context, id string) string {
out, err := winexec.Run(ctx,
"$ErrorActionPreference='SilentlyContinue'\n"+
"(Get-Service -Name "+psQuote(id)+").DisplayName")
if err != nil {
return ""
}
return trimLine(out)
}
@@ -0,0 +1,42 @@
package workloads
import (
"context"
"os"
"time"
"gitea.hostxtra.co.uk/mrhid6/vantage/agent/internal/winexec"
)
const servicesTimeout = 60 * time.Second
const servicesScript = `
$ErrorActionPreference = 'Stop'
$svcs = Get-CimInstance Win32_Service |
Select-Object Name,DisplayName,State,StartMode,PathName,ExitCode
ConvertTo-Json -InputObject @($svcs) -Depth 3 -Compress
`
// collectUnits enumerates Windows services. The bool and string it returns are
// the same SystemdOK / SystemdError pair the Linux collector fills: the wire
// shape is shared, and the UI words it per platform.
func collectUnits(ctx context.Context) ([]Workload, bool, string) {
ctx, cancel := context.WithTimeout(ctx, servicesTimeout)
defer cancel()
out, err := winexec.Run(ctx, servicesScript)
if err != nil {
return nil, false, "Win32_Service query failed: " + err.Error()
}
systemRoot := os.Getenv("SystemRoot")
if systemRoot == "" {
systemRoot = `C:\Windows`
}
wls, err := parseServices(out, systemRoot)
if err != nil {
return nil, false, "Win32_Service output could not be read: " + err.Error()
}
return wls, true, ""
}
@@ -14,10 +14,10 @@ const systemdTimeout = 30 * time.Second
// ten anyone cares about.
var excludedPrefixes = []string{"systemd-", "user@", "user-", "session-", "init.scope"}
// collectSystemd enumerates services in two passes, because "running or
// collectUnits enumerates services in two passes, because "running or
// failed" and "enabled but stopped" are different questions — and an enabled
// unit that is not running is exactly the one worth seeing.
func collectSystemd(ctx context.Context) ([]Workload, bool, string) {
func collectUnits(ctx context.Context) ([]Workload, bool, string) {
if _, err := exec.LookPath("systemctl"); err != nil {
return nil, false, ""
}
+22
View File
@@ -0,0 +1,22 @@
//go:build !linux && !windows
// The build constraint is load-bearing — see updates_other.go.
package workloads
import (
"context"
"fmt"
)
var ownContainerID = ""
func collectUnits(context.Context) ([]Workload, bool, string) { return nil, false, "" }
func isProtectedUnit(string, string) bool { return false }
func controlPlatform(context.Context, string, string, string) error {
return fmt.Errorf("workload control is not supported on this platform")
}
func logsPlatform(context.Context, string, string, int) (string, error) {
return "", fmt.Errorf("workload logs are not supported on this platform")
}
+216
View File
@@ -0,0 +1,216 @@
package workloads
import (
"encoding/json"
"strings"
)
// winService is one row of Get-CimInstance Win32_Service.
//
// Win32_Service rather than Get-Service: Get-Service exposes neither PathName
// nor StartMode, and the filter below needs both.
type winService struct {
Name string `json:"Name"`
DisplayName string `json:"DisplayName"`
State string `json:"State"`
StartMode string `json:"StartMode"`
PathName string `json:"PathName"`
ExitCode int `json:"ExitCode"`
}
// exitCodeNeverStarted is ERROR_SERVICE_NEVER_STARTED. A stopped service
// carrying it has not failed — it has not run since boot — and painting that
// red would cry wolf on every host.
const exitCodeNeverStarted = 1077
// servicePath extracts the executable from a Win32_Service PathName.
//
// A naive split on whitespace misfiles a substantial share of a real fleet:
// `"C:\Program Files\X\x.exe" -service` is one path and one argument.
func servicePath(pathName string) string {
s := strings.TrimSpace(pathName)
if s == "" {
return ""
}
if s[0] == '"' {
if end := strings.IndexByte(s[1:], '"'); end >= 0 {
return s[1 : 1+end]
}
// No closing quote: a malformed or truncated PathName. Fall back to
// the unquoted handling below on the text after the opening quote,
// so this yields a bare path rather than a path plus trailing
// argument text.
s = s[1:]
}
if i := exeBoundaryIndex(s); i >= 0 {
return s[:i+len(".exe")]
}
if i := strings.IndexAny(s, " \t"); i >= 0 {
return s[:i]
}
return s
}
// exeBoundaryIndex finds the first ".exe" (case-insensitive) in s that
// actually ends the executable name — followed by end-of-string, whitespace,
// or a double quote — rather than continuing into a longer segment such as
// ".exec". It returns -1 when no such occurrence exists, so a path like
// `C:\Program Files\Ad.exec\tool.com -flag` is not misparsed by matching the
// ".exe" inside "Ad.exec" and silently dropping the real filename.
func exeBoundaryIndex(s string) int {
lower := strings.ToLower(s)
from := 0
for {
rel := strings.Index(lower[from:], ".exe")
if rel < 0 {
return -1
}
idx := from + rel
end := idx + len(".exe")
if end == len(s) || s[end] == ' ' || s[end] == '\t' || s[end] == '"' {
return idx
}
from = idx + 1
}
}
// parseServices turns the collector's JSON into workloads.
//
// systemRoot is a parameter rather than an environment read so this is testable
// off Windows. The caller passes %SystemRoot%.
//
// The filter mirrors the systemd collector's intent: show what an operator
// installed, and show what is meant to be up but is not. Services under
// %SystemRoot%\System32 are the platform's own, and a typical host has well
// over a hundred of them.
func parseServices(jsonText, systemRoot string) ([]Workload, error) {
s := strings.TrimSpace(jsonText)
if s == "" || s == "null" {
return nil, nil
}
var rows []winService
if err := json.Unmarshal([]byte(s), &rows); err != nil {
var one winService
if err2 := json.Unmarshal([]byte(s), &one); err2 != nil {
return nil, err
}
rows = []winService{one}
}
sys32 := strings.ToLower(strings.TrimRight(systemRoot, `\`) + `\system32\`)
var wls []Workload
for _, r := range rows {
if p := strings.ToLower(servicePath(r.PathName)); p != "" && strings.HasPrefix(p, sys32) {
continue
}
running := strings.EqualFold(r.State, "Running")
failed := !running && r.ExitCode != 0 && r.ExitCode != exitCodeNeverStarted
auto := strings.HasPrefix(strings.ToLower(r.StartMode), "auto")
if !running && !failed && !auto {
continue
}
state := "stopped"
switch {
case running:
state = "running"
case failed:
state = "failed"
}
name := r.DisplayName
if name == "" {
name = r.Name
}
wls = append(wls, Workload{
Kind: "unit",
ID: r.Name,
Name: name,
State: state,
})
}
return wls, nil
}
// psQuote renders a Go string as a PowerShell single-quoted literal. Single
// quotes suppress every form of expansion, so the only character needing an
// escape is the quote itself, which is doubled.
func psQuote(s string) string {
return "'" + strings.ReplaceAll(s, "'", "''") + "'"
}
// scmProvider is the provider every service's start and stop is logged under,
// host-wide.
const scmProvider = "Service Control Manager"
type winEvent struct {
T string `json:"t"`
L string `json:"l"`
P string `json:"p"`
M string `json:"m"`
}
// parseEvents renders Get-WinEvent output as text in the shape journalctl
// --output=short-iso produces, so the log dialog needs no per-platform
// rendering: "<timestamp> <level> <message>", oldest first.
//
// The caller over-fetches from Get-WinEvent because the ProviderName filter
// includes the host-wide Service Control Manager, and a -MaxEvents cap
// applied before SCM rows are narrowed down to this service would squeeze the
// target's own events out of the window on a host with busy service churn.
// tail is therefore applied here, AFTER filtering and AFTER the oldest-first
// reversal, keeping the last tail lines — the most recent lines are the ones
// worth keeping, matching capLog's front-trim reasoning in the shared
// logs.go.
func parseEvents(jsonText, serviceName, displayName string, tail int) (string, error) {
s := strings.TrimSpace(jsonText)
if s == "" || s == "null" {
return "", nil
}
var rows []winEvent
if err := json.Unmarshal([]byte(s), &rows); err != nil {
var one winEvent
if err2 := json.Unmarshal([]byte(s), &one); err2 != nil {
return "", err
}
rows = []winEvent{one}
}
var lines []string
for _, e := range rows {
if strings.EqualFold(e.P, scmProvider) {
if !strings.Contains(e.M, serviceName) &&
(displayName == "" || !strings.Contains(e.M, displayName)) {
continue
}
}
msg := strings.TrimSpace(strings.ReplaceAll(e.M, "\r\n", " "))
lines = append(lines, e.T+" "+e.L+" "+msg)
}
// Get-WinEvent is newest-first. Reverse it.
for i, j := 0, len(lines)-1; i < j; i, j = i+1, j-1 {
lines[i], lines[j] = lines[j], lines[i]
}
if tail > 0 && len(lines) > tail {
lines = lines[len(lines)-tail:]
}
return strings.Join(lines, "\n"), nil
}
// trimLine reduces single-value PowerShell output to its first non-empty line.
func trimLine(s string) string {
for _, l := range strings.Split(s, "\n") {
if t := strings.TrimSpace(l); t != "" {
return t
}
}
return ""
}
+188
View File
@@ -0,0 +1,188 @@
package workloads
import (
"strings"
"testing"
)
func TestServicePath(t *testing.T) {
cases := []struct{ in, want string }{
{`"C:\Program Files\Contoso\svc.exe" -service`, `C:\Program Files\Contoso\svc.exe`},
{`C:\WINDOWS\system32\svchost.exe -k netsvcs`, `C:\WINDOWS\system32\svchost.exe`},
{`C:\Vantage\vantage-agent.exe`, `C:\Vantage\vantage-agent.exe`},
{`"C:\no\args.exe"`, `C:\no\args.exe`},
{``, ``},
// ".exe" appearing inside an earlier segment ("Ad.exec") must not be
// treated as the end of the executable — that would drop the real
// filename and arguments.
{`C:\Program Files\Ad.exec\tool.com -flag`, `C:\Program`},
// An unterminated quote falls back to the unquoted handling on the
// text after the opening quote, yielding a bare path rather than a
// path plus trailing argument text.
{`"C:\Program Files\Contoso\svc.exe -service`, `C:\Program Files\Contoso\svc.exe`},
}
for _, c := range cases {
if got := servicePath(c.in); got != c.want {
t.Errorf("servicePath(%q) = %q, want %q", c.in, got, c.want)
}
}
}
func TestParseServicesFilters(t *testing.T) {
in := `[
{"Name":"Contoso","DisplayName":"Contoso Broker","State":"Running","StartMode":"Auto","PathName":"\"C:\\Program Files\\Contoso\\svc.exe\" -service","ExitCode":0},
{"Name":"Themes","DisplayName":"Themes","State":"Running","StartMode":"Auto","PathName":"C:\\WINDOWS\\system32\\svchost.exe -k netsvcs","ExitCode":0},
{"Name":"Fabrikam","DisplayName":"Fabrikam Sync","State":"Stopped","StartMode":"Auto","PathName":"C:\\Fabrikam\\sync.exe","ExitCode":0},
{"Name":"Northwind","DisplayName":"Northwind Poller","State":"Stopped","StartMode":"Manual","PathName":"C:\\Northwind\\poll.exe","ExitCode":0},
{"Name":"Crashed","DisplayName":"Crashed Thing","State":"Stopped","StartMode":"Auto","PathName":"C:\\Crashed\\c.exe","ExitCode":1067}
]`
got, err := parseServices(in, `C:\WINDOWS`)
if err != nil {
t.Fatalf("parseServices: %v", err)
}
byID := map[string]Workload{}
for _, w := range got {
byID[w.ID] = w
}
// The OS's own svchost service is dropped; a manual, stopped, never-failed
// service is nobody's business either.
if _, ok := byID["Themes"]; ok {
t.Error("Themes (under %SystemRoot%) should be filtered out")
}
if _, ok := byID["Northwind"]; ok {
t.Error("stopped Manual service should be filtered out")
}
if len(got) != 3 {
t.Fatalf("got %d workloads, want 3: %+v", len(got), got)
}
if w := byID["Contoso"]; w.Kind != "unit" || w.Name != "Contoso Broker" || w.State != "running" {
t.Errorf("Contoso = %+v", w)
}
// Enabled but not running is exactly the row worth seeing.
if byID["Fabrikam"].State != "stopped" {
t.Errorf("Fabrikam state = %q, want stopped", byID["Fabrikam"].State)
}
// A non-zero exit code on a stopped service is a crash, not a clean stop.
if byID["Crashed"].State != "failed" {
t.Errorf("Crashed state = %q, want failed", byID["Crashed"].State)
}
}
// 1077 means "no attempt to start since boot" — a clean stopped service, not a
// failure, and reporting it red would cry wolf on every host.
func TestParseServicesExitCode1077(t *testing.T) {
in := `[{"Name":"Idle","DisplayName":"Idle","State":"Stopped","StartMode":"Auto","PathName":"C:\\Idle\\i.exe","ExitCode":1077}]`
got, err := parseServices(in, `C:\WINDOWS`)
if err != nil {
t.Fatalf("parseServices: %v", err)
}
if len(got) != 1 || got[0].State != "stopped" {
t.Fatalf("got %+v, want one stopped workload", got)
}
}
func TestParseServicesSingleObjectAndEmpty(t *testing.T) {
one := `{"Name":"Solo","DisplayName":"Solo","State":"Running","StartMode":"Auto","PathName":"C:\\Solo\\s.exe","ExitCode":0}`
got, err := parseServices(one, `C:\WINDOWS`)
if err != nil || len(got) != 1 {
t.Fatalf("single object: got %+v, err %v", got, err)
}
for _, in := range []string{"", "[]", "null"} {
got, err := parseServices(in, `C:\WINDOWS`)
if err != nil || len(got) != 0 {
t.Fatalf("parseServices(%q) = %+v, err %v", in, got, err)
}
}
}
func TestPSQuote(t *testing.T) {
if got := psQuote(`it's`); got != `'it''s'` {
t.Fatalf("psQuote = %s", got)
}
if got := psQuote(`plain`); got != `'plain'` {
t.Fatalf("psQuote = %s", got)
}
}
func TestParseEventsFormatsAndOrders(t *testing.T) {
// Get-WinEvent returns newest first; journalctl --output=short-iso returns
// oldest first, and the log dialog and capLog's front-trim both assume the
// most recent line is at the bottom.
in := `[
{"t":"2026-08-13T10:22:31.0000000Z","l":"Error","p":"Contoso","m":"broker died"},
{"t":"2026-08-13T10:22:03.0000000Z","l":"Information","p":"Contoso","m":"broker starting"}
]`
got, err := parseEvents(in, "Contoso", "Contoso Broker", 500)
if err != nil {
t.Fatalf("parseEvents: %v", err)
}
want := "2026-08-13T10:22:03.0000000Z Information broker starting\n" +
"2026-08-13T10:22:31.0000000Z Error broker died"
if got != want {
t.Fatalf("parseEvents =\n%q\nwant\n%q", got, want)
}
}
// Service Control Manager logs every service on the host under one provider, so
// its rows must be filtered down to the target or the log is somebody else's.
func TestParseEventsFiltersOtherServicesSCM(t *testing.T) {
in := `[
{"t":"2026-08-13T10:00:00Z","l":"Information","p":"Service Control Manager","m":"The Print Spooler service entered the running state."},
{"t":"2026-08-13T10:00:01Z","l":"Information","p":"Service Control Manager","m":"The Contoso Broker service entered the running state."}
]`
got, err := parseEvents(in, "Contoso", "Contoso Broker", 500)
if err != nil {
t.Fatalf("parseEvents: %v", err)
}
if strings.Contains(got, "Print Spooler") {
t.Errorf("another service's SCM event leaked in:\n%s", got)
}
if !strings.Contains(got, "Contoso Broker") {
t.Errorf("the target's SCM event was dropped:\n%s", got)
}
}
// A service that has logged nothing is normal. An error there would read as a
// broken feature.
func TestParseEventsEmpty(t *testing.T) {
for _, in := range []string{"", "[]", "null"} {
got, err := parseEvents(in, "Contoso", "Contoso Broker", 500)
if err != nil || got != "" {
t.Fatalf("parseEvents(%q) = %q, err %v", in, got, err)
}
}
}
// The over-fetch in logs_windows.go can return more events than the caller
// asked for once SCM rows are filtered down to the target; parseEvents must
// keep the most RECENT tail lines, not the oldest, matching capLog's
// front-trim reasoning in the shared logs.go.
func TestParseEventsTrimsToTailKeepingMostRecent(t *testing.T) {
in := `[
{"t":"2026-08-13T10:00:06Z","l":"Information","p":"Contoso","m":"event 6"},
{"t":"2026-08-13T10:00:05Z","l":"Information","p":"Contoso","m":"event 5"},
{"t":"2026-08-13T10:00:04Z","l":"Information","p":"Contoso","m":"event 4"},
{"t":"2026-08-13T10:00:03Z","l":"Information","p":"Contoso","m":"event 3"},
{"t":"2026-08-13T10:00:02Z","l":"Information","p":"Contoso","m":"event 2"},
{"t":"2026-08-13T10:00:01Z","l":"Information","p":"Contoso","m":"event 1"}
]`
got, err := parseEvents(in, "Contoso", "Contoso Broker", 2)
if err != nil {
t.Fatalf("parseEvents: %v", err)
}
want := "2026-08-13T10:00:05Z Information event 5\n" +
"2026-08-13T10:00:06Z Information event 6"
if got != want {
t.Fatalf("parseEvents =\n%q\nwant\n%q", got, want)
}
}
+3 -7
View File
@@ -4,7 +4,6 @@ import (
"context"
"crypto/sha256"
"encoding/hex"
"runtime"
"sort"
"strconv"
"strings"
@@ -19,15 +18,12 @@ type Result struct {
SystemdError string
}
// Collect enumerates every workload on this host. Linux only.
// Collect enumerates every workload on this host: containers from Docker, and
// units from systemd on Linux or the service control manager on Windows.
func Collect(ctx context.Context) Result {
if runtime.GOOS != "linux" {
return Result{}
}
var r Result
containers, dockerOK, dockerErr := collectDocker(ctx)
units, systemdOK, systemdErr := collectSystemd(ctx)
units, systemdOK, systemdErr := collectUnits(ctx)
r.DockerOK, r.DockerError = dockerOK, dockerErr
r.SystemdOK, r.SystemdError = systemdOK, systemdErr
-30
View File
@@ -1,30 +0,0 @@
# Current cloud instance process
The current processs for creating cloud instances is incorrect.
At the moment the cloud instance process is the following:
- Customer goes to `https://vantage.hostxtra.co.uk/start` then fills in the form.
- Customer is then sent and email to verify
- Customer clicks the link and the instance is created in the DB.
- Customer can then access the instance.
As the `/start` process is auto creating a new instance this should default to the free tier instance.
The issue is that this doesn't create an `account` and `admin_instance` on the admin side.
The Cloud instance creation / account creation needs to be restructured.
for context when I say `hq` I mean `admin`
- customer goes to `https://vantage.hostxtra.co.uk/start` and fills in the form.
- This is where the HQ account is created.
- The customer is then sent and email to verify their email address.
- The customer can then access the HQ customer portal.
- The customer can then create a free new instance in the HQ portal.
- The cloud instance is created in the DB.
- The HQ `account` and `admin_instance` is created and populated in the DB.
- A `Free` License is created and attached to the instance.
- The customer is then sent an email letting them know the instance has been created and when the license expires.
- The customer will need to renew the license after expiry, if they are on a Free license.
- This is so that unused instances can be cleaned up if no renew after a length of time has passed.
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -1,938 +0,0 @@
# Instance Rename in Vantage HQ — Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Let an HQ customer (owner or admin) rename a cloud instance, which re-derives its slug and moves it to a new DNS host, with staff able to do the same without the cooldown.
**Architecture:** Slug derivation stays in `shared/provision`, beside the create path that already owns it. Admin reaches the control plane only through `cloudprov`, writing `instances` — a collection it already writes. Admin's own row (`admin_instances`) is updated second and carries the 24h cooldown timestamp, because the cooldown is admin's policy and the control plane has no opinion about it. The portal shows the new host and asks the customer to click through; it does not redirect.
**Note:** This repo has no automated test suite and the user has ruled out adding test files. Every task verifies by build, vet and (Task 8) manual exercise.
**Tech Stack:** Go 1.x (gin, mongo-driver v2), Next.js 16 App Router + TanStack Query + Tailwind 3 (`adminsite`).
**Spec:** `docs/superpowers/specs/2026-08-12-instance-rename-design.md`
## Global Constraints
- A licence binds an instance **UUID**, not a slug. A rename must not issue a licence, call Paddle, or touch `licenses`, `subscriptions` or `entitlements`.
- Admin's control-plane write boundary is unchanged: `cloudprov` writes `instances` and `users` only. Do not add a write to any other control-plane collection.
- Customer rename is **cloud only**. Self-hosted is refused with the existing `selfHostedRefusal` constant and HTTP **400**, matching `members.go`.
- Cooldown for customers is **24 hours**, tracked by `admin_instances.renamed_at`. Staff bypass it and must **not** write `renamed_at`.
- No `-2` suffix loop on rename. A taken slug is a refusal (`ErrSlugTaken` → HTTP 409).
- No component in `adminsite` may carry a hex colour; use the existing token classes (`text-ink-2`, `text-ink-3`, `border-rule`, `text-accent`, `text-expired`, `bg-panel-2`).
- The host domain used for display is `vantage.hostxtra.co.uk`, already hardcoded in `adminsite/components/InstanceRecord.tsx` and the customer instance page.
- Commit messages follow the repo's existing style: `feat: Sentence case summary` / `fix: …` / `docs: …`.
---
### Task 1: Slug derivation and the control-plane rename
**Files:**
- Modify: `shared/provision/instance.go`
**Interfaces:**
- Consumes: `BaseSlug(name string) (string, error)`, `ErrNameRejected` — both already in `shared/provision`.
- Produces:
- `provision.ErrSlugTaken` (`error`)
- `provision.RenameSlug(name, currentSlug string) (string, error)`
- `provision.RenameInstance(ctx context.Context, db *mongo.Database, instanceID, name string) (*models.Instance, error)`
- `provision.RestoreInstanceIdentity(ctx context.Context, db *mongo.Database, instanceID, name, slug string) error`
Behaviour `RenameSlug` must have, verified by reading rather than by test (this
repo has no Go test suite and the user has ruled out adding one):
| Input name | Current slug | Result |
|---|---|---|
| `Acme Ltd` | `acme` | `acme-ltd` |
| `ACME!` | `acme` | `acme` — still derives to the current slug, so not a move |
| `Acme` | `acme-2` | `acme` — a creation-time collision suffix derives from no name, so moving off it is a real move |
| `ab` | any | `ErrNameRejected` |
| `Admin` | any | `ErrNameRejected` (reserved) |
| `!!!` | any | `ErrNameRejected` |
| 50 `a`s | any | truncated to `MaxSlugLength`, exactly as `BaseSlug` truncates on create |
- [ ] **Step 1: Write the implementation**
Append to `shared/provision/instance.go`:
```go
// ErrSlugTaken means the slug a new name derives to already belongs to another
// instance.
//
// Rename refuses rather than appending a counter the way creation does. Creation
// appends because the customer is waiting on an instance and any free slug will
// do; a rename is a request for one specific host, and silently landing them on
// "acme-2" answers a question they did not ask.
var ErrSlugTaken = errors.New("slug taken")
// RenameSlug derives the slug a rename to name would move an instance to, given
// the slug it holds now.
//
// It returns the current slug unchanged when the name still derives to it, so a
// cosmetic edit — capitalisation, punctuation, a trailing "Ltd." — is not a move
// and cannot collide with the instance's own slug.
func RenameSlug(name, currentSlug string) (string, error) {
base, err := BaseSlug(name)
if err != nil {
return "", fmt.Errorf("%w: %s", ErrNameRejected, err.Error())
}
if base == currentSlug {
return currentSlug, nil
}
return base, nil
}
// RenameInstance changes an instance's name and re-derives its slug from it.
//
// The count-then-update is racy on its own, and is safe for the same reason
// CreateInstanceWithID's loop is: instances.slug carries a unique index, so a
// lost race surfaces as a duplicate-key error. Unlike creation there is nothing
// to retry with — the caller asked for one specific name — so it becomes
// ErrSlugTaken. Do not remove the duplicate-key branch, and do not remove the
// index.
func RenameInstance(ctx context.Context, db *mongo.Database, instanceID, name string) (*models.Instance, error) {
var inst models.Instance
if err := db.Collection("instances").FindOne(ctx,
bson.M{"instance_id": instanceID}).Decode(&inst); err != nil {
return nil, err
}
slug, err := RenameSlug(name, inst.Slug)
if err != nil {
return nil, err
}
if slug != inst.Slug {
n, err := db.Collection("instances").CountDocuments(ctx, bson.M{
"slug": slug,
"instance_id": bson.M{"$ne": instanceID},
})
if err != nil {
return nil, err
}
if n > 0 {
return nil, fmt.Errorf("%w: %s", ErrSlugTaken, slug)
}
}
if _, err := db.Collection("instances").UpdateOne(ctx,
bson.M{"instance_id": instanceID},
bson.M{"$set": bson.M{"name": name, "slug": slug}}); err != nil {
if mongo.IsDuplicateKeyError(err) {
return nil, fmt.Errorf("%w: %s", ErrSlugTaken, slug)
}
return nil, err
}
inst.Name = name
inst.Slug = slug
return &inst, nil
}
// RestoreInstanceIdentity writes an exact name and slug back, unwinding a rename
// whose caller-side bookkeeping then failed.
//
// It derives nothing. The values being restored may include a creation-time
// collision suffix that no name derives to, so re-running RenameInstance with the
// old name would not reproduce them.
func RestoreInstanceIdentity(ctx context.Context, db *mongo.Database, instanceID, name, slug string) error {
_, err := db.Collection("instances").UpdateOne(ctx,
bson.M{"instance_id": instanceID},
bson.M{"$set": bson.M{"name": name, "slug": slug}})
return err
}
```
- [ ] **Step 2: Build and vet**
Run: `cd /go-projects/vantage && go build ./shared/... && go vet ./shared/provision/`
Expected: clean.
- [ ] **Step 3: Commit**
```bash
git add shared/provision/instance.go
git commit -m "feat: Add instance rename to shared provisioning"
```
---
### Task 2: Admin's row and the cloudprov wrappers
**Files:**
- Modify: `admin/internal/models/models.go` (the `Instance` struct, ~line 129; constants block near `RenewWindow`, ~line 105)
- Modify: `admin/internal/cloudprov/cloudprov.go`
**Interfaces:**
- Consumes: `provision.RenameInstance`, `provision.RestoreInstanceIdentity` (Task 1).
- Produces:
- `models.RenameCooldown` (`time.Duration`)
- `models.Instance.RenamedAt *time.Time` (bson `renamed_at`, json `renamed_at`)
- `cloudprov.RenameInstance(ctx context.Context, instanceID, name string) (*sharedmodels.Instance, error)`
- `cloudprov.RestoreInstanceIdentity(ctx context.Context, instanceID, name, slug string) error`
- [ ] **Step 1: Add the cooldown constant**
In `admin/internal/models/models.go`, directly beneath the `RenewWindow` block:
```go
// RenameCooldown is how long a customer must wait between renames of one
// instance.
//
// A rename moves the instance's DNS host and invalidates every saved link to it,
// so this exists to make that a considered act rather than a slider. Staff are
// not subject to it: a support conversation about a name is already a human
// deciding.
const RenameCooldown = 24 * time.Hour
```
- [ ] **Step 2: Add the field to `Instance`**
In the same file, inside the `Instance` struct, after `RelinkCount`:
```go
// RenamedAt is when this instance last changed name, and backs the customer
// rename cooldown. It is a pointer because absent means "never renamed"; a
// zero time.Time would read as year 1 — an inert cooldown, but only by
// accident. Staff renames deliberately leave it alone.
RenamedAt *time.Time `bson:"renamed_at,omitempty" json:"renamed_at,omitempty"`
```
- [ ] **Step 3: Add the cloudprov wrappers**
Append to `admin/internal/cloudprov/cloudprov.go`:
```go
// RenameInstance changes a cloud instance's name and moves it to the slug that
// name derives to.
//
// It writes `instances` and nothing else, so admin's control-plane write
// boundary is unchanged. It issues no licence: a licence binds the instance
// UUID, which a rename never touches.
func RenameInstance(ctx context.Context, instanceID, name string) (*sharedmodels.Instance, error) {
return provision.RenameInstance(ctx, db.ControlDB(), instanceID, name)
}
// RestoreInstanceIdentity puts an instance's previous name and slug back, for a
// caller unwinding a rename whose admin-side write failed. Leaving the two
// databases disagreeing would have HQ print a host that is not the host.
func RestoreInstanceIdentity(ctx context.Context, instanceID, name, slug string) error {
return provision.RestoreInstanceIdentity(ctx, db.ControlDB(), instanceID, name, slug)
}
```
- [ ] **Step 4: Build**
Run: `cd /go-projects/vantage && go build ./admin/... ./shared/...`
Expected: clean build, no output.
- [ ] **Step 5: Commit**
```bash
git add admin/internal/models/models.go admin/internal/cloudprov/cloudprov.go
git commit -m "feat: Add rename cooldown field and cloudprov rename"
```
---
### Task 3: Customer rename endpoint
**Files:**
- Modify: `admin/internal/api/customer.go` (add handler; `loginURLFor` at ~line 442 is already in this file)
- Modify: `admin/internal/api/routes.go` (~line 77, beside the other `/instances/:id/*` customer routes)
**Interfaces:**
- Consumes: `ownedInstance(c, id) (*models.Instance, bool)`, `selfHostedRefusal` (`members.go`), `loginURLFor(slug) string`, `cloudprov.RenameInstance`, `cloudprov.RestoreInstanceIdentity`, `models.RenameCooldown`, `provision.ErrSlugTaken`, `provision.ErrNameRejected`.
- Produces: `PUT /api/instances/:id/name` returning `{instance_id, name, slug, login_url}`.
- [ ] **Step 1: Write the handler**
Append to `admin/internal/api/customer.go`:
```go
// renameInstance changes a cloud instance's name and moves it to the slug that
// name derives to.
//
// The control plane is written FIRST, because instances.slug carries the unique
// index and that index is what actually settles a race between two accounts
// reaching for the same name. Admin's own row follows; if that write fails the
// control plane is put back, because HQ printing a host that is not the host is
// worse than a failed rename.
//
// No licence is issued and Paddle is not called: a licence binds the instance
// UUID, and a rename does not change it.
func renameInstance(c *gin.Context) {
inst, ok := ownedInstance(c, c.Param("id"))
if !ok {
return
}
if inst.Deployment != license.DeploymentCloud {
c.JSON(http.StatusBadRequest, gin.H{"error": selfHostedRefusal})
return
}
if inst.Placeholder {
c.JSON(http.StatusConflict, gin.H{"error": "this instance is not provisioned yet"})
return
}
var body struct {
Name string `json:"name"`
}
if err := c.ShouldBindJSON(&body); err != nil {
c.JSON(http.StatusBadRequest, gin.H{"error": "name is required"})
return
}
name := strings.TrimSpace(body.Name)
if name == "" {
c.JSON(http.StatusBadRequest, gin.H{"error": "name is required"})
return
}
if inst.RenamedAt != nil {
if until := inst.RenamedAt.Add(models.RenameCooldown); time.Now().UTC().Before(until) {
c.JSON(http.StatusTooManyRequests, gin.H{
"error": fmt.Sprintf("this instance was renamed recently; it can be renamed again after %s UTC", until.Format("2 Jan 2006 15:04")),
"retry_after": until,
})
return
}
}
ctx := c.Request.Context()
renamed, err := cloudprov.RenameInstance(ctx, inst.InstanceID, name)
switch {
case errors.Is(err, provision.ErrSlugTaken):
c.JSON(http.StatusConflict, gin.H{"error": "that name is already in use — try another"})
return
case errors.Is(err, provision.ErrNameRejected):
c.JSON(http.StatusUnprocessableEntity, gin.H{"error": err.Error()})
return
case err != nil:
log.Printf("renameInstance: control plane rename of %s: %v", inst.InstanceID, err)
c.JSON(http.StatusInternalServerError, gin.H{"error": "could not rename the instance"})
return
}
if _, err := db.Admin("admin_instances").UpdateOne(ctx,
bson.M{"instance_id": inst.InstanceID},
bson.M{"$set": bson.M{
"name": renamed.Name,
"slug": renamed.Slug,
"renamed_at": time.Now().UTC(),
}}); err != nil {
if rbErr := cloudprov.RestoreInstanceIdentity(ctx, inst.InstanceID, inst.Name, inst.Slug); rbErr != nil {
log.Printf("renameInstance: rollback of %s failed: %v", inst.InstanceID, rbErr)
}
log.Printf("renameInstance: record rename of %s: %v", inst.InstanceID, err)
c.JSON(http.StatusInternalServerError, gin.H{"error": "could not rename the instance"})
return
}
s := auth.Current(c)
audit.Write(ctx, models.AuditEntry{
Actor: s.Email, Action: "instance.renamed", AccountID: s.AccountID,
Target: inst.InstanceID, Detail: inst.Slug + " -> " + renamed.Slug, IP: c.ClientIP()})
c.JSON(http.StatusOK, gin.H{
"instance_id": inst.InstanceID,
"name": renamed.Name,
"slug": renamed.Slug,
// The same builder the licence emails use, rather than a second opinion
// about how a tenant host is spelled. Empty when APP_LOGIN_URL is unset.
"login_url": loginURLFor(renamed.Slug),
})
}
```
- [ ] **Step 2: Check the imports**
`customer.go` must import `errors`, `fmt`, `log`, `net/http`, `strings`, `time`, `audit`, `auth`, `cloudprov`, `db`, `models`, `license`, `provision`, `gin`, `bson`. Most are already there — add only what the compiler asks for. `provision` is `gitea.hostxtra.co.uk/mrhid6/vantage/shared/provision`; `license` is `gitea.hostxtra.co.uk/mrhid6/vantage/shared/license`.
- [ ] **Step 3: Mount the route**
In `admin/internal/api/routes.go`, in the `cust` group beside the other instance routes (after `cust.POST("/instances/:id/claim-free", …)`):
```go
// Renaming moves the instance's DNS host, so it is owner-or-admin like
// every other instance mutation. Cloud only; the handler refuses the rest.
cust.PUT("/instances/:id/name",
auth.RequireAccountRole(models.AccountRoleOwner, models.AccountRoleAdmin),
renameInstance)
```
- [ ] **Step 4: Build**
Run: `cd /go-projects/vantage && go build ./admin/... && go vet ./admin/internal/api/`
Expected: clean.
- [ ] **Step 5: Commit**
```bash
git add admin/internal/api/customer.go admin/internal/api/routes.go
git commit -m "feat: Add customer instance rename endpoint"
```
---
### Task 4: Staff rename endpoint
**Files:**
- Modify: `admin/internal/api/staff.go`
- Modify: `admin/internal/api/routes.go` (the `staff` group, beside `staff.POST("/instances/:id/relink", …)`)
**Interfaces:**
- Consumes: everything Task 3 consumes, plus `db.Admin`.
- Produces: `PUT /api/staff/instances/:id/name` returning `{instance_id, name, slug}`.
- [ ] **Step 1: Write the handler**
Append to `admin/internal/api/staff.go`:
```go
// staffRenameInstance renames any instance, with no cooldown.
//
// It does NOT write renamed_at: a staff rename must not start the customer's
// 24h clock, or fixing a name for someone locks them out of fixing it further.
//
// On self-hosted it changes admin's label only. There is no control-plane row to
// write — the install is the customer's — and no slug, because self-hosted has
// no tenant subdomain.
func staffRenameInstance(c *gin.Context) {
var body struct {
Name string `json:"name"`
}
if err := c.ShouldBindJSON(&body); err != nil {
c.JSON(http.StatusBadRequest, gin.H{"error": "name is required"})
return
}
name := strings.TrimSpace(body.Name)
if name == "" {
c.JSON(http.StatusBadRequest, gin.H{"error": "name is required"})
return
}
ctx := c.Request.Context()
var inst models.Instance
if err := db.Admin("admin_instances").FindOne(ctx,
bson.M{"instance_id": c.Param("id")}).Decode(&inst); err != nil {
c.JSON(http.StatusNotFound, gin.H{"error": "not found"})
return
}
set := bson.M{"name": name}
slug := inst.Slug
if inst.Deployment == license.DeploymentCloud && !inst.Placeholder {
renamed, err := cloudprov.RenameInstance(ctx, inst.InstanceID, name)
switch {
case errors.Is(err, provision.ErrSlugTaken):
c.JSON(http.StatusConflict, gin.H{"error": "that name is already in use"})
return
case errors.Is(err, provision.ErrNameRejected):
c.JSON(http.StatusUnprocessableEntity, gin.H{"error": err.Error()})
return
case err != nil:
c.JSON(http.StatusInternalServerError, gin.H{"error": err.Error()})
return
}
slug = renamed.Slug
set["slug"] = renamed.Slug
}
if _, err := db.Admin("admin_instances").UpdateOne(ctx,
bson.M{"instance_id": inst.InstanceID}, bson.M{"$set": set}); err != nil {
if inst.Deployment == license.DeploymentCloud && !inst.Placeholder {
if rbErr := cloudprov.RestoreInstanceIdentity(ctx, inst.InstanceID, inst.Name, inst.Slug); rbErr != nil {
log.Printf("staffRenameInstance: rollback of %s failed: %v", inst.InstanceID, rbErr)
}
}
c.JSON(http.StatusInternalServerError, gin.H{"error": err.Error()})
return
}
audit.Write(ctx, models.AuditEntry{
Actor: auth.Current(c).Email, Action: "instance.renamed", AccountID: inst.AccountID,
Target: inst.InstanceID, Detail: inst.Slug + " -> " + slug, IP: c.ClientIP()})
c.JSON(http.StatusOK, gin.H{"instance_id": inst.InstanceID, "name": name, "slug": slug})
}
```
`staff.go` will need `errors`, `log`, `cloudprov` and `provision` added to its imports; `fmt`, `net/http`, `strings`, `time`, `audit`, `auth`, `db`, `models`, `license`, `bson` are already there.
- [ ] **Step 2: Mount the route**
In `routes.go`, in the `staff` group after `staff.POST("/instances/:id/relink", staffRelink)`:
```go
staff.PUT("/instances/:id/name", staffRenameInstance)
```
- [ ] **Step 3: Build**
Run: `cd /go-projects/vantage && go build ./admin/... && go vet ./admin/internal/api/`
Expected: clean.
- [ ] **Step 4: Commit**
```bash
git add admin/internal/api/staff.go admin/internal/api/routes.go
git commit -m "feat: Add staff instance rename endpoint"
```
---
### Task 5: `adminsite` API client and slug preview
**Files:**
- Create: `adminsite/lib/slug.ts`
- Modify: `adminsite/lib/api.ts` (the `Instance` interface ~line 123; the `api` object's instance calls ~line 305; `api.staff` ~line 360)
**Interfaces:**
- Consumes: `PUT /api/instances/:id/name`, `PUT /api/staff/instances/:id/name` (Tasks 3 and 4).
- Produces:
- `INSTANCE_DOMAIN`, `slugify(name: string): string`, `slugError(name: string): string | undefined` from `@/lib/slug`
- `RenameResult` interface, `api.renameInstance(id, name): Promise<RenameResult>`, `api.staff.renameInstance(id, name): Promise<RenameResult>`
- `Instance.renamed_at?: string`
- [ ] **Step 1: Create the slug mirror**
Create `adminsite/lib/slug.ts`:
```ts
/*
* A TypeScript mirror of shared/provision's slug rules, used ONLY to preview the
* host a rename would move an instance to while the customer types.
*
* It is a second implementation of Slugify, BaseSlug and ReservedSlugs, and it
* must change in the same commit as the Go one — the same hazard as
* web/lib/targets.ts. The preview is a courtesy; the server's 409 is the
* boundary, and the two are allowed to disagree without anything breaking.
*/
/** Mirrors provision.MinSlugLength / MaxSlugLength. */
export const MIN_SLUG_LENGTH = 3;
export const MAX_SLUG_LENGTH = 40;
/** Mirrors provision.ReservedSlugs. */
const RESERVED = new Set([
"www", "api", "app", "admin", "auth",
"install", "static", "_next", "default",
]);
/*
* The tenant subdomain namespace. Also hardcoded in InstanceRecord.tsx and the
* customer instance page; those predate this file and are left alone rather than
* refactored under a rename change.
*/
export const INSTANCE_DOMAIN = "vantage.hostxtra.co.uk";
/** Mirrors provision.Slugify. */
export function slugify(name: string): string {
return name
.toLowerCase()
.replace(/[^a-z0-9]+/g, "-")
.replace(/^-+|-+$/g, "");
}
/** Mirrors provision.BaseSlug's truncation. */
export function baseSlug(name: string): string {
return slugify(name).slice(0, MAX_SLUG_LENGTH);
}
/** The reason a name cannot become a slug, or undefined when it can. */
export function slugError(name: string): string | undefined {
const base = slugify(name);
if (base.length < MIN_SLUG_LENGTH) {
return `Needs at least ${MIN_SLUG_LENGTH} letters or digits.`;
}
if (RESERVED.has(base.slice(0, MAX_SLUG_LENGTH))) {
return "That name is reserved.";
}
return undefined;
}
/** The host an instance on this slug is reached at. */
export function hostFor(slug: string): string {
return `${slug}.${INSTANCE_DOMAIN}`;
}
```
- [ ] **Step 2: Extend the API client**
In `adminsite/lib/api.ts`, add `renamed_at` to `Instance` (after `relink_count`):
```ts
renamed_at?: string;
```
Add the response type beside the other interfaces:
```ts
export interface RenameResult {
instance_id: string;
name: string;
slug: string;
/** Empty when APP_LOGIN_URL is unset on the server. */
login_url?: string;
}
```
Add the call to the `api` object, after `renewInstance`:
```ts
renameInstance: (id: string, name: string) =>
put<RenameResult>(`/api/instances/${id}/name`, { name }),
```
And to `api.staff`, after `relink`:
```ts
renameInstance: (id: string, name: string) =>
put<RenameResult>(`/api/staff/instances/${id}/name`, { name }),
```
- [ ] **Step 3: Type-check**
Run: `cd /go-projects/vantage/adminsite && npx tsc --noEmit`
Expected: no errors.
- [ ] **Step 4: Commit**
```bash
git add adminsite/lib/slug.ts adminsite/lib/api.ts
git commit -m "feat: Add rename calls and slug preview to the HQ client"
```
---
### Task 6: The rename panel and the customer instance page
**Files:**
- Create: `adminsite/components/RenamePanel.tsx`
- Modify: `adminsite/app/(customer)/instances/[id]/page.tsx`
**Interfaces:**
- Consumes: `api.renameInstance` / `api.staff.renameInstance`, `RenameResult` (Task 5); `slugError`, `baseSlug`, `hostFor` (Task 5); `Panel`, `Note` (`@/components/Panel`), `Button` (`@/components/Button`), `Field` (`@/components/Field`), `ApiError` (`@/lib/api`).
- Produces: `RenamePanel({ currentName, currentSlug, onRename })` — a default-collapsed control; `onRename` is `(name: string) => Promise<RenameResult>`.
- [ ] **Step 1: Create the component**
Create `adminsite/components/RenamePanel.tsx`:
```tsx
"use client";
import { useState } from "react";
import { Button } from "./Button";
import { Field } from "./Field";
import { Note } from "./Panel";
import { ApiError, type RenameResult } from "@/lib/api";
import { baseSlug, hostFor, slugError } from "@/lib/slug";
/*
* The rename control, and only the control — the same shape as RelinkPanel: an
* input that expands in place rather than a modal, because this app has no modal
* and one action with one field does not need one.
*
* The host preview is drawn from lib/slug.ts, a mirror of the Go rules. It can
* disagree with the server; the 409 that comes back is the answer that counts.
*/
export function RenamePanel({
currentName,
currentSlug,
onRename,
}: {
currentName: string;
currentSlug: string;
onRename: (name: string) => Promise<RenameResult>;
}) {
const [open, setOpen] = useState(false);
const [value, setValue] = useState(currentName);
const [error, setError] = useState<string | undefined>();
const [busy, setBusy] = useState(false);
const [done, setDone] = useState<RenameResult | undefined>();
const name = value.trim();
const derived = baseSlug(name);
const invalid = slugError(name);
// A cosmetic edit that lands on the same slug is still a rename worth doing —
// the name is what the customer reads. Only an empty or unchanged name is
// nothing to submit.
const unchanged = name === currentName.trim();
async function submit() {
setError(undefined);
setBusy(true);
try {
const res = await onRename(name);
setDone(res);
setOpen(false);
} catch (err) {
setError(err instanceof ApiError ? err.message : "Rename failed. Try again.");
} finally {
setBusy(false);
}
}
if (done) {
const host = done.login_url || `https://${hostFor(done.slug)}`;
return (
<Note tone="warn">
<span className="grid gap-2">
<span>
This instance is now <strong>{done.name}</strong>, at{" "}
<span className="font-mono">{hostFor(done.slug)}</span>. The old address has stopped working, and your
sign-in does not follow it — you will need to sign in again there.
</span>
<a href={host} className="justify-self-start font-mono text-[0.78rem] text-accent underline">
Open {hostFor(done.slug)} &rarr;
</a>
</span>
</Note>
);
}
return (
<div className="grid gap-3">
{open && (
<Field
label="Instance name"
value={value}
onChange={(e) => setValue(e.target.value)}
error={error ?? (name ? invalid : undefined)}
hint={
name && !invalid ? (
<>
Moves to <span className="font-mono">{hostFor(derived)}</span>
{derived === currentSlug && " — the address does not change"}
</>
) : (
"Letters and digits; everything else becomes a hyphen."
)
}
/>
)}
<div className="flex flex-wrap items-center gap-3">
<Button
type="button"
variant="line"
disabled={busy || (open && (!name || Boolean(invalid) || unchanged))}
onClick={() => (open ? submit() : setOpen(true))}
>
{busy ? "Renaming…" : "Rename instance"}
</Button>
{open && (
<span className="text-[0.82rem] text-ink-3">
Anyone signed in will need to sign in again at the new address, and links to the old one stop working.
</span>
)}
</div>
</div>
);
}
```
`Note` is `({ tone = "accent" | "warn" | "expired", children })` and renders a `<p>`, which is why the success state wraps its two lines in a `<span className="grid gap-2">` rather than block elements.
- [ ] **Step 2: Mount it on the customer instance page**
In `adminsite/app/(customer)/instances/[id]/page.tsx`:
Add the imports:
```tsx
import { RenamePanel } from "@/components/RenamePanel";
```
and
```tsx
import { useSession } from "@/lib/session";
```
Inside `InstancePage`, with the other hooks (hooks must precede the early returns already in this component):
```tsx
// useSession is the app's one way to ask who the caller is — it shares the
// ["me"] query, so this adds no request.
const { session } = useSession();
```
and after the `cloud` const:
```tsx
const mayRename = session?.account_role === "owner" || session?.account_role === "admin";
```
Then add the panel to `PageFrame`'s children, directly after the `MembersPanel` line:
```tsx
{/*
* Address rather than "Rename": the panel is about where this
* instance lives, and the rename is how you change it. Cloud
* only — a self-hosted install has no tenant subdomain for us to
* move.
*/}
{cloud && mayRename && (
<Panel title="Address" meta={host ?? undefined}>
<p className="text-[0.86rem] text-ink-2">
The instance name is where its address comes from. Renaming moves it to a new address and releases the old
one, so saved links and bookmarks to it stop working.
</p>
<RenamePanel
currentName={instance.name}
currentSlug={instance.slug ?? ""}
onRename={async (name) => {
const res = await api.renameInstance(instance.instance_id, name);
qc.invalidateQueries({ queryKey: ["account"] });
return res;
}}
/>
</Panel>
)}
```
- [ ] **Step 3: Build**
Run: `cd /go-projects/vantage/adminsite && npm run build`
Expected: build succeeds.
- [ ] **Step 4: Commit**
```bash
git add adminsite/components/RenamePanel.tsx "adminsite/app/(customer)/instances/[id]/page.tsx"
git commit -m "feat: Let a customer rename a cloud instance from HQ"
```
---
### Task 7: Staff instance page rename
**Files:**
- Modify: `adminsite/app/(staff)/staff/instances/[id]/page.tsx`
**Interfaces:**
- Consumes: `RenamePanel` (Task 6), `api.staff.renameInstance` (Task 5).
- Produces: nothing later tasks depend on.
- [ ] **Step 1: Add the panel**
In `adminsite/app/(staff)/staff/instances/[id]/page.tsx`, add the imports:
```tsx
import { RenamePanel } from "@/components/RenamePanel";
```
and, inside `StaffInstancePage`, add `const qc = useQueryClient();` at the top of the component if it is not already there (`useQueryClient` is already imported for `EntitlementSection`).
Add this panel after the "Licence history" panel:
```tsx
{/*
* Staff rename has no cooldown and does not start the customer's:
* fixing a name on someone's behalf must not spend their next 24
* hours.
*/}
<Panel title="Name" meta={data.instance.deployment === "cloud" ? "Moves the address" : "Label only"}>
<RenamePanel
currentName={data.instance.name}
currentSlug={data.instance.slug ?? ""}
onRename={async (name) => {
const res = await api.staff.renameInstance(data.instance.instance_id, name);
qc.invalidateQueries({ queryKey: ["staff-instance", id] });
return res;
}}
/>
</Panel>
```
- [ ] **Step 2: Build**
Run: `cd /go-projects/vantage/adminsite && npm run build`
Expected: build succeeds.
- [ ] **Step 3: Commit**
```bash
git add "adminsite/app/(staff)/staff/instances/[id]/page.tsx"
git commit -m "feat: Let staff rename an instance"
```
---
### Task 8: Documentation and end-to-end verification
**Files:**
- Modify: `CLAUDE.md` (the Admin REST API route list, and the `admin_instances` note under MongoDB Collections)
**Interfaces:**
- Consumes: everything above.
- Produces: nothing.
- [ ] **Step 1: Update the Admin REST API route list**
In `CLAUDE.md`, in the customer-session block, after the `POST /instances/:id/claim-free` line:
```
PUT /instances/:id/name # rename a cloud instance; moves its slug (owner|admin, 24h cooldown)
```
and in the staff-session block, after `POST /instances/:id/issue · /instances/:id/relink`:
```
PUT /instances/:id/name # rename any instance, no cooldown
```
- [ ] **Step 2: Add the design note**
In `CLAUDE.md`, under "Grants project, they do not federate" (admin's control-plane write boundary is described nearby), add a short paragraph:
```markdown
**A rename moves the host, and the licence does not care.** `PUT
/api/instances/:id/name` re-derives the slug from the new name through
`provision.RenameSlug` — the same rules that named the instance at creation —
and writes the control plane first, because `instances.slug`'s unique index is
what settles a race between two accounts reaching for one name. A taken slug is
a refusal, not an `acme-2`: creation appends a counter because any free slug
will do, and a rename is a request for one specific host. A licence binds the
instance UUID, so nothing is reissued and Paddle is not called. The old host
keeps resolving for up to 60s (`instancehost.go`'s cache, which admin cannot
reach into), and `km_session` is host-only, so the customer signs in again on
the new address — the portal says so rather than redirecting them into a login
screen with no explanation. The 24h cooldown lives on `admin_instances.renamed_at`
because it is admin's policy; staff bypass it and must not write the field.
```
- [ ] **Step 3: Full build**
Run:
```bash
cd /go-projects/vantage && go build ./... && go vet ./admin/... ./shared/... && (cd adminsite && npm run build)
```
Expected: all clean.
- [ ] **Step 4: Manual verification against a running stack**
Work through each and record the result:
1. Rename a cloud instance from `/instances/<id>` as an owner. Panel reports the new host.
2. In Mongo: `db.instances.findOne({instance_id})` and `db.admin_instances.findOne({instance_id})` agree on `name` and `slug`; `admin_instances.renamed_at` is set.
3. The new host serves a login page. The old host stops resolving to the instance within ~60 seconds.
4. A second rename inside 24 hours answers `429` with the unlock time.
5. Renaming onto a slug another instance holds answers `409` and changes neither database.
6. `PUT /api/instances/:id/name` on a self-hosted instance answers `400` with the `selfHostedRefusal` message.
7. `GET /api/staff/audit` shows `instance.renamed` with `old-slug -> new-slug`.
8. Staff rename of the same instance succeeds immediately and leaves `renamed_at` unchanged.
- [ ] **Step 5: Commit**
```bash
git add CLAUDE.md
git commit -m "docs: Document instance rename in HQ"
```
- [ ] **Step 6: Refresh the knowledge graph**
```bash
graphify update .
```
File diff suppressed because it is too large Load Diff
@@ -1,307 +0,0 @@
# Multiple auth providers
Date: 2026-08-03
## Problem
An instance can configure exactly one OIDC provider. `instance_oidc` holds one
document per instance, `/auth/oidc/start` takes no argument, and `/login`
renders an unconditional "Sign in with your instance's SSO" button whether or
not anything is configured behind it. Customers who federate with more than one
identity source cannot, and customers who federate with none are shown a button
that leads to an error.
## Goals
- N auth providers per instance, each independently enabled and named.
- Login page renders one button per enabled provider, and none when there are
none.
- Local email/password login can be turned off per instance.
- Presets for the common identity providers, so a customer supplies a tenant ID
rather than an issuer URL.
- Existing configured SSO keeps working across the upgrade with no customer
action.
## Non-goals
- SAML. Different protocol, metadata parsing and certificate handling; not in
this work.
- Per-provider role or group mapping. Provisioned users remain `member`, as
today.
- Provider-specific account linking. An email address is an email address; the
existing instance-scoped lookup stands.
## Data model
New collection `auth_providers`, one document per provider:
```go
type AuthProvider struct {
ID bson.ObjectID `bson:"_id,omitempty" json:"_id,omitempty"`
InstanceID string `bson:"instance_id" json:"instance_id"`
ProviderID string `bson:"provider_id" json:"provider_id"`
Name string `bson:"name" json:"name"`
Kind string `bson:"kind" json:"kind"` // "oidc" | "oauth2"
Preset string `bson:"preset" json:"preset"` // "" for custom
Issuer string `bson:"issuer" json:"issuer"`
ClientID string `bson:"client_id" json:"client_id"`
ClientSecretEnc string `bson:"client_secret_enc,omitempty" json:"-"`
Scopes []string `bson:"scopes" json:"scopes"`
Enabled bool `bson:"enabled" json:"enabled"`
CallbackNotice bool `bson:"callback_notice" json:"callback_notice"`
Order int `bson:"order" json:"order"`
CreatedAt time.Time `bson:"created_at" json:"created_at"`
UpdatedAt time.Time `bson:"updated_at" json:"updated_at"`
}
```
`ProviderID` is a short random identifier, not the Mongo `_id`: it appears in
the callback URL a customer pastes into their IdP, and an `_id` there would
publish a database key.
Unique index on `(instance_id, provider_id)`. Index build is fatal on failure,
matching `EnsureAuthIndexes` — a duplicate `provider_id` within an instance
would make the callback ambiguous.
`ClientSecretEnc` is AES-256-GCM under `KEY_ENCRYPTION_KEY`, as
`instance_oidc.client_secret_enc` is today, and is never serialised.
### Presets
A Go table in `server/internal/auth/presets.go`, not database rows — adding one
is a commit, not a migration.
| Preset | Kind | Issuer | Input asked of the customer | Default scopes |
| ------- | -------- | ------------------------------------------------- | --------------------------- | ----------------------------- |
| `entra` | `oidc` | `https://login.microsoftonline.com/{tenant}/v2.0` | Directory (tenant) ID | `openid profile email` |
| `google`| `oidc` | `https://accounts.google.com` | none | `openid profile email` |
| `okta` | `oidc` | `https://{domain}/oauth2/default` | Okta org domain | `openid profile email` |
| `github`| `oauth2` | n/a | none | `read:user user:email` |
| `` (custom) | `oidc` | supplied verbatim | Issuer URL | `openid profile email` |
The issuer template is expanded server-side on save; the stored `Issuer` is
always the resolved URL, so nothing downstream has to know a preset existed.
### Settings
`settings.local_login_enabled bool`, defaulting true. Absent on existing
documents, and Go's zero value for `bool` is false, so the field is read through
a `*bool` and a nil pointer means enabled. A plain `bool` would silently
disable password login on every instance in the fleet at upgrade.
## Migration
`0005_auth_providers` — the next free number; `0004_instance_rename` is the
highest recorded today. For each document in `instance_oidc`, insert one
`auth_providers` document:
- `Name: "Single sign-on"`
- `Preset: ""`, `Kind: "oidc"`
- `Issuer`, `ClientID`, `Enabled` copied
- `ClientSecretEnc` copied **verbatim**, not decrypted and re-encrypted — a
migration that needs `KEY_ENCRYPTION_KEY` fails on an instance that has none
and strands the SSO configuration.
- `Scopes: ["openid", "profile", "email"]`, matching what `oidc.go` hardcodes
today.
- `ProviderID` freshly generated.
- `CallbackNotice: true` — this provider's redirect URI has changed and an
administrator has not yet acknowledged it. Set only by the migration; cleared
by the settings UI. New providers are created `false`.
`instance_oidc` is left in place and no longer read. Idempotent by skipping any
instance that already has an `auth_providers` document, so a re-run after a
partial failure completes rather than duplicating.
## Auth flow
Routes:
```
GET /auth/oidc/:providerId/start
GET /auth/oidc/:providerId/callback
```
The old unparameterised `/auth/oidc/start` and `/auth/oidc/callback` are
**removed**, not retained. See Upgrade impact below — this breaks configured SSO
until the customer updates their IdP, and that is accepted deliberately rather
than carried as a compatibility path.
The state token in Redis stores `{instance_id, provider_id}` rather than the
bare instance ID. The callback resolves its provider from the consumed state
and cross-checks it against `:providerId` in the path, refusing a mismatch —
the path alone is attacker-controlled, and the state is the half that was
issued by the start handler.
`providerForInstance` becomes `providerFor(ctx, c, instanceID, providerID)`.
The `go-oidc` provider cache keys on `provider_id`, not instance. Saving,
disabling or deleting a provider evicts that key.
`redirectURL(c, providerID)` returns the one per-provider shape, and returns the
same URL in the start and callback halves of a flow — an IdP rejects the token
exchange if they differ.
### OIDC providers
Unchanged from the current implementation: `AuthCodeURL` with the stored
scopes, exchange, `id_token` verified against the provider's key set with
`ClientID` as audience, `email` and `name` claims extracted.
### GitHub (`kind: "oauth2"`)
GitHub is OAuth2 and issues no `id_token`, so it takes a separate branch:
exchange the code, then `GET https://api.github.com/user/emails` with the access
token and take the address that is both `primary` and `verified`. An
unverified-only response is refused — an unverified address is not proof of
control, and accepting one would let anyone holding a GitHub account claim any
address in the instance. `name` comes from `GET https://api.github.com/user`.
Both branches converge on one function:
```go
func completeSSOLogin(c *gin.Context, instanceID, email, name string) error
```
which holds today's lookup-or-provision, session creation, `TouchLastLogin` and
cookie set, verbatim. Email is lower-cased before lookup, and the lookup stays
`GetUserInInstanceByEmail` — instance-scoped, as it is now.
### Licence gate
`services.GetLicenseState(instanceID).Feature("oidc")` continues to gate both
the start and the callback, for every provider kind, and is checked on the
callback against the instance named by the consumed state rather than the host.
Unchanged behaviour, applied to more providers.
## REST API
Unauthenticated:
```
GET /auth/providers
-> {"local_enabled": true,
"providers": [{"id": "...", "name": "...", "preset": "entra"}]}
```
Instance is resolved from the host, as `/auth/bootstrap-status` already does.
The response carries **no issuer, no client ID and no secret** — it is served to
anyone who can reach the login page.
Session-authed, `owner|admin`, under `/api`:
```
GET,POST /auth/providers
PUT,DELETE /auth/providers/:id
POST /auth/providers/:id/test
```
`test` fetches the provider's discovery document (or, for GitHub, calls the API
with the stored credentials) and reports reachability. It does not sign anyone
in.
`GET,PUT /api/org/oidc` is removed along with the old auth routes. Its only
caller is `OIDCCard.tsx`, which this work replaces, and a compatibility shim
over a one-of-many model would have to invent which provider it means.
Every mutation writes an audit event, as every mutating path does.
### Lockout guards
Both refused with 409 and a distinct error code:
- `local_login_required` — disabling local login while zero providers are
enabled.
- `last_provider` — disabling or deleting the last enabled provider while local
login is off.
These are enforced in the service layer, not the handler, so the two endpoints
that can reach the condition cannot disagree.
## Frontend
### Settings
`web/components/settings/OIDCCard.tsx` becomes `AuthProvidersCard`, in the
Access group of `/settings` where the OIDC card already lives. It renders the
provider list with per-row enable toggle, edit, delete and drag ordering, an
Add flow that asks for the preset first and then only the fields that preset
needs, and the local-login toggle beneath the list. A guard violation surfaces
the 409's message rather than a generic failure.
Every provider row shows its **callback URL** with click-to-copy — that is the
value the customer pastes into their IdP, it now differs per provider, and after
the upgrade every migrated provider needs it re-pasted. A migrated provider
additionally carries a warning until an administrator dismisses it, naming the
change and the URL. Dismissal is per provider, stored on the document.
### Login page
`web/app/login/page.tsx` calls `/auth/providers` on mount alongside the existing
`bootstrapStatus` call, and renders on the result:
| `local_enabled` | providers | Rendered |
| --------------- | --------- | --------------------------------------------------- |
| true | none | Password form only. No divider, no buttons. |
| true | some | Password form, divider, one button per provider. |
| false | some | Buttons only. No form, no divider. |
| false | none | Password form (see below). |
The last row cannot be reached through the API — the guards above prevent it —
but a hand-edited database could produce it, and a login page that renders
nothing at all is unrecoverable without database access. It therefore falls back
to the password form.
The current unconditional SSO button and its "SSO must be enabled for this
instance by an administrator" note are both removed; the button now only exists
when it works.
Buttons are labelled with the provider's `Name` and carry the preset's icon
where there is one, a neutral key glyph otherwise. Presets never override the
name — a customer who calls their Entra provider "Staff" gets "Staff".
Errors keep the existing `/login?error=<code>` redirect convention.
## Testing
- Migration: an `instance_oidc` document produces one enabled provider with the
ciphertext byte-identical; a re-run inserts nothing further.
- `local_login_enabled` absent decodes as enabled.
- Guards: both 409 paths, and the enable/disable sequences that approach them
without crossing.
- Per-provider callback: two providers in one instance, each resolving to its
own configuration; a `provider_id` from another instance answers 404.
- A callback whose `:providerId` disagrees with the consumed state is refused,
and the state is consumed rather than left replayable.
- The removed routes (`/auth/oidc/start`, `/auth/oidc/callback`,
`/api/org/oidc`) answer 404.
- GitHub: primary+verified selected; verified-only-absent refused.
- `/auth/providers` response contains no issuer, client ID or secret.
## Upgrade impact
**This release breaks configured SSO until each customer updates their identity
provider.** The old `/auth/oidc/callback` is gone, migrated providers are
reachable only at `/auth/oidc/<providerId>/callback`, and an IdP still pointing
at the old URL fails the flow.
It is a deliberate trade: one callback shape rather than two, no
`legacy_callback` branch through `redirectURL`, and no permanently retained
route whose only purpose is a single past upgrade.
Mitigations, in order of who sees them first:
- The settings card shows the new callback URL per provider with click-to-copy,
and a migrated provider carries a dismissable warning naming the change.
- The failure is visible rather than silent: an IdP rejects the redirect URI
before Vantage is reached, so the customer sees their own provider's error.
- Local password login is unaffected, so no instance is locked out — an
administrator can always sign in to fix the URL. This is why
`local_login_enabled` defaults to true and why nothing in this migration
turns it off.
- Release notes and `docsite/docs/vantage/settings.md` state the required
action.
## Deployment notes
No new environment variables. No agent change. `KEY_ENCRYPTION_KEY` is already
required wherever OIDC was configured, and the migration does not add a
dependency on it.
@@ -1,235 +0,0 @@
# Server tags and scheduled workflows
Date: 2026-08-04
Two features, designed together because the second is worth much less without
the first. Tags make a target set describable; schedules make it recur. A
nightly job that patches "everything tagged `env:staging`" needs both halves,
and neither half is large on its own.
---
## Part A — Server tags
### Model
`models.Server` gains one field:
```go
Tags map[string]string `bson:"tags,omitempty" json:"tags,omitempty"`
```
Keys and values are lowercase `[a-z0-9_-]`. Keys are capped at 32 characters,
values at 64, and a server holds at most 20 tags. Validation lives in the
service layer rather than the handler, so the tag endpoint, the server-create
path and anything added later cannot disagree about what a valid tag is.
There is **no `tags` collection.** A tag is a property of a server, not an
entity with a lifecycle: a registry would need reference counting to know when
a tag stopped existing, and garbage collection to act on it, which is work
bought for nothing. The list of known keys and values that the UI offers for
autocomplete is a distinct aggregation over `servers`, cached for 60 seconds —
the same treatment org lookups already get.
No reserved keys ship in this change. If inventory-derived tags (`os`, `arch`)
are added later they take a `sys:` key prefix, so a user tag written today can
never collide with a system tag invented tomorrow.
Index: `{instance_id: 1, "tags.$**": 1}` — a wildcard index over the tag
subdocument, because the queried key is chosen by the user at request time and
cannot be named in advance.
### API
```
PUT /api/servers/:id/tags # replace the whole map
GET /api/servers/tags # known keys and values, for pickers
GET /api/servers?tag=env:prod # repeatable; AND across keys
```
`PUT` replaces the entire map rather than patching one tag. A tag set is small
enough that sending all of it is free, and last-write-wins over a whole map is
easier to reason about than merge semantics between two people editing the same
server. The audit event records the map before and after.
`?tag=` is repeatable and ANDs: `?tag=env:prod&tag=role:web` matches servers
carrying both. A malformed value (no colon, unknown characters) is a 400 rather
than a silent empty result — a filter that matches nothing and a filter that is
nonsense look identical in a list, and only one of them is the user's fault.
### Targeting
`models.Workflow` gains `TargetTags map[string]string` beside the existing
`TargetServerIDs`. One function in `services` resolves them:
```go
ResolveTargets(ctx, instanceID string, ids []string, tags map[string]string) ([]Server, error)
```
- Result is the **distinct union** of the explicit IDs and the tag matches.
- Tag matching ANDs across keys.
- Offline servers are included. The dispatcher already answers 503 per server,
and a patch run that silently omits an unreachable machine is worse than one
that visibly fails on it.
- Empty IDs **and** empty tags returns `ErrNoTargets` (400). A workflow that
matches nothing must say so rather than report success over zero servers.
The resolved set is snapshotted into `WorkflowRun.ServerRuns` exactly as today.
History records what actually ran, not what the selector would match when the
run is later read back — the same reason `steps_snapshot` exists.
### Frontend
- **Server detail**: tag chips in the header with an inline editor. Keys
autocomplete from `GET /api/servers/tags`, values autocomplete per key.
- **`/servers`**: a filter bar that reads and writes the same `?tag=` query
params the API takes, so a filtered fleet view is a URL someone can send.
- **Workflow designer**: a target section holding both inputs, with a live
"runs on 14 servers" readout that lists them on hover. The union model costs
us the at-a-glance answer to "what will this touch"; this readout buys it
back, and it is the reason the union is acceptable.
---
## Part B — Scheduled workflows
### Model
```go
type Schedule struct {
Enabled bool `bson:"enabled" json:"enabled"`
Cron string `bson:"cron" json:"cron"` // 5-field
TZ string `bson:"tz" json:"tz"` // IANA name
}
type Skip struct {
Reason string `bson:"reason" json:"reason"` // "missed" | "already_running"
Due time.Time `bson:"due" json:"due"`
At time.Time `bson:"at" json:"at"`
}
```
On `Workflow`:
```go
Schedule *Schedule `bson:"schedule,omitempty"`
NextRunAt *time.Time `bson:"next_run_at,omitempty"` // UTC, indexed
LastRunAt *time.Time `bson:"last_run_at,omitempty"`
LastSkipped *Skip `bson:"last_skipped,omitempty"`
```
`next_run_at` is **persisted, not held in memory.** A leader handover between
computing the next occurrence and firing it would otherwise either lose the
occurrence or fire it twice. Coordination state has to live where every replica
can see it — the same argument that put `workflow_log_seq` in MongoDB.
Cron parsing uses `robfig/cron/v3`'s **parser only**`Parse` and
`Next(time)`. Its scheduler and goroutines are not used; the loop below is ours
and has to be, because it runs under the leader lock.
**Alpine ships no tzdata.** `server/Dockerfile` builds a slim image, so
`time.LoadLocation("Europe/London")` returns an error and every schedule
falls back to UTC — an hour wrong for half the year, in the direction nobody
notices until a maintenance window lands in business hours. `main` therefore
imports `_ "time/tzdata"`, embedding the database in the binary. Zone names are
also validated at save time, so an unknown zone is a 400 rather than a surprise
at 2am.
### Scheduler
A new `server/internal/workflowsched` package, started inside the **existing**
`bus.RunAsLeader("housekeeping", …)` alongside `monitorsched`, `StartReaper`
and the sweepers. One role, one lock. It takes the same cancellable context and
returns the instant leadership is lost.
The loop ticks every 30 seconds:
1. `find({schedule.enabled: true, next_run_at: {$lte: now}})`.
2. **Claim atomically.** `findOneAndUpdate` matching the document *and* its
current `next_run_at`, setting the recomputed next occurrence. A process
that reaches the same document after another has claimed it matches nothing
and does nothing. The claim is what makes this correct; the leader lock only
makes it cheap.
3. **Grace check.** If `now - due > 1h`, record
`last_skipped{reason: "missed"}`, write an audit event, and do not run. A
job missed by ten minutes during a deploy should still run; one missed by
two days should not fire at lunchtime.
4. **Overlap check.** If a run for this workflow is still active, record
`last_skipped{reason: "already_running"}`, audit, and do not run. A patch
workflow must never run twice at once, and a silent skip is how a week goes
by before anyone notices nothing ran.
5. Otherwise start the run through the **same** `RunWorkflow` path a person
uses, with `TriggeredBy: "schedule"`.
Step 5 is the design. A scheduled run is an ordinary run with a different
trigger: no second dispatch path, no second snapshot format, and the run detail
page needs no changes to display one.
### API
```
PUT /api/workflows/:id/schedule # {enabled, cron, tz}
GET /api/workflows/:id/schedule/preview?cron=…&tz=… # next 3 occurrences
```
`PUT` validates the expression and the zone, then computes and stores
`next_run_at`. The preview endpoint exists so the browser and the scheduler
agree on what a cron string means — a client-side cron parser that disagrees
with the server by one field is a bug found in production, at night.
### Frontend
- **Workflow page**: a schedule card with preset buttons (hourly, nightly at
HH:MM, weekly on DAY at HH:MM) that write cron underneath, a raw cron field
for anything else, a timezone select, and the next three occurrences rendered
from the preview endpoint in mono.
- **Workflows list**: a schedule chip and the next run as relative time.
- **Skips are surfaced**, not just stored: a warning line reading
"Skipped Sun 02:00 — previous run still active". Recording a reason nobody
reads is the same as not recording one.
---
## Out of scope
**Notification on scheduled-run failure.** It needs the monitor channel
machinery pointed at workflow outcomes and its own answer to what counts as
failure — a non-zero exit on a step with `on_failure: continue` is not
obviously an alert. Visibility in this change is the run list and the recorded
skip reason. Excluded deliberately, not overlooked.
**Tag-scoped permissions.** Roles stay instance-wide. Tags describe servers;
they do not yet gate who may act on them.
**Inventory-derived tags.** Reserved via the `sys:` prefix, not implemented.
---
## Migration and compatibility
No migration is required. `Tags`, `TargetTags` and `Schedule` are all
`omitempty` and absent means what it meant before: no tags, no selector, no
schedule. Existing workflows keep their explicit server lists and behave
identically.
The wildcard tag index and the `next_run_at` index are declared by a new
`EnsureServerIndexes`, following the convention `EnsureSecretIndexes` and
`EnsureWorkflowIndexes` already set: it warns rather than aborting boot,
because a missing index degrades
tag filtering to a collection scan on a small collection rather than breaking
the fleet list.
## Testing
- `ResolveTargets`: union deduplicates; AND across tag keys; empty/empty
returns `ErrNoTargets`; offline servers are included.
- Tag validation: charset, length caps, tag count cap, malformed `?tag=` is a
400.
- Schedule validation: bad cron and unknown zone both 400; `next_run_at` is
computed in the stored zone, verified across a DST boundary.
- Scheduler claim: two concurrent claims of the same due workflow start exactly
one run.
- Grace window: due 10 minutes ago runs; due 2 hours ago records `missed`.
- Overlap: an active run yields `already_running` and no second run.
- Preview endpoint and the scheduler agree on the next occurrence for a table
of expressions, including a DST-crossing one.
@@ -1,538 +0,0 @@
# Package inventory and CVE findings
Date: 2026-08-06
Agents report the packages installed on each server. The control plane matches
them against distro security feeds and raises findings that link straight to
the patching path that already exists. A finding nobody can fix today can be
accepted with a reason and an expiry date rather than sitting red forever.
This is one of four sub-projects sketched together and deliberately separated:
| # | Sub-project | Depends on |
| - | ----------- | ---------- |
| A | **Package inventory + CVE findings** — this spec | nothing |
| B | Container/service registry | nothing |
| C | Container image scanning | A and B |
| D | Compliance profiles (baseline assertions) | shares A's findings UI only |
A and B are independent of one another. C is the joiner and must not be
designed before both exist. D shares a page with A and nothing else — a
different collector, a different evaluation model and a different remediation
story — so folding it in here would double the size for no shared machinery.
Scope of this spec is **A, Linux only.** Windows needs a separate source
(MSRC CVRF), a separate collector (`Get-HotFix` plus registry) and a KB
supersedence matcher that shares no code with the Linux path. That matches the
existing position that Windows agents are second-class by design, and the six
package managers `updates.go` already detects cover the whole Linux surface.
---
## The trap this design is built around
Distributions **backport** security fixes without changing the upstream
version. Ubuntu ships `openssl 3.0.2-0ubuntu1.15` patched against
CVE-2023-0286; NVD says version 3.0.2 is vulnerable. Matching installed
versions against NVD or CPE ranges therefore reports a fleet full of criticals
that are all already fixed.
That is not merely noisy. It is fatal to the feature: once the first report is
mostly wrong, nobody reads the second one, and a genuine finding is lost in the
noise it created. Everything below follows from refusing to make that mistake.
The correct source is the **distribution's own security feed**, keyed on the
distribution's own version string — Debian and Ubuntu OVAL/USN, Red Hat OVAL
v2, Alpine secdb. `trivy-db` is those feeds pre-merged into one BoltDB
artifact, rebuilt every six hours and published as an OCI artifact.
---
## Where the vulnerability data comes from
`trivy-db`, pulled server-side from `ghcr.io/aquasecurity/trivy-db:2`.
The alternative considered was querying OSV.dev per scan, which needs no
storage and no puller. It was rejected on two counts: it requires outbound
internet on every scan, which breaks air-gapped installs; and it sends the
package list of a customer's entire fleet to a third party. The audience most
likely to buy vulnerability scanning is the audience least willing to do that.
The blob is roughly 50MB, read-only, reproducible, and identified by a version
number. **It is not stored in Mongo and not written to `/data`**
`server.persistence` defaults to off and nothing writes to `/data` any more.
It does not need durable storage: whichever pod needs it pulls it to its own
ephemeral temp directory. Nothing shared, nothing to back up, nothing to
migrate.
`VANTAGE_TRIVY_DB_REF` overrides the default reference so a customer can mirror
the artifact into their own registry. It also covers the anonymous ghcr rate
limit, which the six-hourly pull cadence already makes unlikely to bite.
---
## Only the leader matches
This is the crux, and it falls out of the replica model already in the
codebase.
Two things trigger matching, and they happen on different pods:
1. a fleet-wide rescan when `trivy-db` updates — naturally the leader's job
2. a server's package list changing — handled by whichever pod holds *that
agent's* command stream
If (2) matched inline, **every replica would need the 50MB database resident**,
and a database refresh would have N pods racing to rescan the same fleet and N
digests reaching the customer. That is the exact failure `RunAsLeader` exists
to prevent, and it is the same argument that put `monitorsched` behind the
lock.
So `ReportPackages` does not match. It upserts the package list and sets
`scan_pending: true`. That is all it does.
`server/internal/vulnsched` then runs inside the **existing**
`bus.RunAsLeader("housekeeping", …)` alongside `monitorsched`,
`workflowsched` and the sweepers — one role, one lock. Every 60 seconds it:
1. pulls `trivy-db` if the local copy is older than six hours
2. if the pulled version differs from `vulndb_meta.db_version`, marks **every**
server `scan_pending`
3. matches all `scan_pending` servers, clears the flag, diffs against existing
findings
4. emits **one** digest per tick covering everything newly opened
Step 4 is why batching is structural rather than bolted on. A `trivy-db`
refresh can open several hundred findings across a fleet at once; one message
per finding would rate-limit the webhook or get the channel muted, and either
way the customer stops receiving the alerts they are paying for. The tick is
already the natural batch boundary, so **the failure cannot occur by
construction** rather than by a debounce someone has to maintain.
`scan_pending` lives on the document rather than in memory, for the same reason
`next_run_at` and `workflow_log_seq` do: a leader handover between marking and
scanning would otherwise lose it. A handover costs the new leader one re-pull
of the database.
The cost of this indirection is up to 60 seconds between an agent reporting a
changed package set and its findings updating. For vulnerability data that is
nothing, and it buys a single matching path instead of two.
---
## Components
```
agent/internal/packages/ collect installed packages + /etc/os-release
proto/ ReportPackages RPC
server/internal/vulndb/ puller, BoltDB access, matcher
server/internal/vulnsched/ leader-owned tick: pull, scan, digest
server/internal/services/ findings, acceptance, alert rules
web/app/(app)/vulnerabilities/ fleet board; plus two server-detail tabs
```
`vulnsched` takes the dependencies it needs — `LogEvent` and the notification
dispatch — as a `vulnsched.Deps` injected from `main.go`, following
`workflowsched`. The manual rescan endpoint does not call into `vulnsched` at
all: it sets `scan_pending` on every server and lets the next tick find them,
so there is no path by which `services` imports the scheduler and no cycle to
avoid later.
---
## The wire path
A new `ReportPackages` RPC on the agent's existing hourly loop — the same
`runUpdateCheck` cadence, reusing `updates.go`'s `detectPM()`.
```protobuf
rpc ReportPackages(ReportPackagesRequest) returns (ReportPackagesResponse);
message ReportPackagesRequest {
string server_id = 1;
string agent_token = 2;
string hash = 3; // sha256 of the sorted list
OSRelease os = 4;
repeated InstalledPackage packages = 5; // omitted when only offering a hash
}
message ReportPackagesResponse {
bool need_full = 1; // hash differs; resend with packages populated
}
```
The agent calls once with `packages` empty. `need_full` true means the hash
differs from what the server holds, and the agent immediately calls again with
the list populated.
The agent sends a SHA-256 of its sorted package list first. If it matches what
the server already holds, the server answers `unchanged` and the ~150KB body is
never sent. A machine's package set changes rarely, so almost every hour costs
one small message, and the rare changed hour costs one extra round trip.
Folding the list into the existing 15-minute `InventoryReport` static snapshot
was rejected: it would re-send ~150KB per server every 15 minutes regardless of
change, roughly 40MB/hour of gRPC traffic on a 100-server fleet to transmit
data that is almost always identical.
---
## Data model
Four new collections. Every one carries `instance_id` except `vulndb_meta`,
which is explained below.
### `server_packages` — one document per server, not per package
```go
type ServerPackages struct {
ID primitive.ObjectID `bson:"_id"`
InstanceID primitive.ObjectID `bson:"instance_id"`
ServerID string `bson:"server_id"`
OS OSRelease `bson:"os"` // family, version_id, arch
Hash string `bson:"hash"` // sha256 of the sorted list
Packages []InstalledPackage `bson:"packages"`
CollectedAt time.Time `bson:"collected_at"`
ScanPending bool `bson:"scan_pending"`
ScannedAt time.Time `bson:"scanned_at"`
Status string `bson:"status"` // ok | unsupported
DBVersion int `bson:"db_version"` // last matched against
}
type InstalledPackage struct {
Name string `bson:"name"`
Version string `bson:"version"` // distro version string, verbatim
Epoch int `bson:"epoch,omitempty"`
Arch string `bson:"arch"`
SourceName string `bson:"source_name,omitempty"`
}
```
One document rather than two thousand is what makes a report a **single atomic
upsert with no delta logic** — the hash already established that something
changed, so there is nothing to reconcile field by field. A typical Linux host
lands near 150KB, comfortably inside the 16MB document limit.
Indexes: `{instance_id, server_id}` unique, and a multikey
`{instance_id, "packages.name"}` for fleet-wide package search.
`SourceName` is not decoration. **Debian and Ubuntu advisories are keyed on the
source package**: a CVE against `openssl` covers the binaries `libssl3`,
`openssl` and `libssl-dev`, so matching on binary name alone misses two of the
three.
`OS.VersionID` selects the feed. Ubuntu 22.04 and 24.04 publish different fixed
versions for the same CVE, so a scan without it is guesswork.
### `vuln_findings` — one document per (server, CVE, package)
```go
type VulnFinding struct {
ID primitive.ObjectID `bson:"_id"`
InstanceID primitive.ObjectID `bson:"instance_id"`
ServerID string `bson:"server_id"`
CVEID string `bson:"cve_id"`
PackageName string `bson:"package_name"`
Installed string `bson:"installed_version"`
FixedIn string `bson:"fixed_in,omitempty"`
Severity string `bson:"severity"`
CVSSScore float64 `bson:"cvss_score,omitempty"`
Title string `bson:"title,omitempty"`
References []string `bson:"references,omitempty"`
State string `bson:"state"` // open | fixed | accepted
FirstSeen time.Time `bson:"first_seen"`
LastSeen time.Time `bson:"last_seen"`
FixedAt *time.Time `bson:"fixed_at,omitempty"`
Accepted *Acceptance `bson:"accepted,omitempty"`
}
type Acceptance struct {
By primitive.ObjectID `bson:"by"`
Reason string `bson:"reason"`
Until time.Time `bson:"until"`
At time.Time `bson:"at"`
}
```
Unique on `{instance_id, server_id, cve_id, package_name}`. That key is what
makes a rescan an idempotent upsert rather than a duplicate factory, and it is
what lets `first_seen` survive across scans. Query index
`{instance_id, state, severity}`.
**An empty `FixedIn` is a real and common state** and must never be conflated
with "not vulnerable". A CVE with no vendor fix published yet is exactly the
finding people most need to see, and also the one that most needs acceptance,
because there is nothing to patch.
Findings are **not deleted when a package is patched**. State moves to `fixed`
with `fixed_at` set, so "what did we remediate last quarter" remains
answerable — which is the question an auditor asks.
### `vulndb_meta` — singleton, deliberately unscoped
`db_version`, `pulled_at`, `last_full_scan_at`, `last_error`. It carries no
`instance_id` because the vulnerability database is a property of the
deployment, not of a tenant. Same reasoning as `migrations`.
### `vuln_alert_rules`
`instance_id`, `name`, `enabled`, `min_severity`, `tags map[string]string`,
`channel_ids []`, timestamps.
The tag filter resolves through **`services.ResolveTargets`**, not a second
matcher. That function is already the single answer to which servers a
selector touches, and an alert rule that disagreed with a workflow about what
`env:prod` means would be worse than having no filter at all.
---
## The matching engine
```
server/internal/vulndb/
pull.go OCI fetch → temp dir, version compare against vulndb_meta
db.go BoltDB open, advisory lookup by (ecosystem, source, version)
match.go per-family matching, severity resolution
version.go dispatch to deb/rpm/apk comparator by OS family
```
Dependencies: `github.com/aquasecurity/trivy-db` for the BoltDB schema, plus
`go-deb-version`, `go-rpm-version` and `go-apk-version` — each a small
standalone module doing one job. The roughly 200 lines of per-distro advisory
lookup are ours.
Importing `trivy` itself was rejected: it would pull a very large transitive
dependency tree into the server binary for one feature, and its Go API carries
no stability guarantee across minor versions. Shelling out to the `trivy`
binary against a generated SBOM was rejected for shipping a second binary in
the image and turning a library call into subprocess lifecycle, timeouts and
output-format drift.
### Why the comparators are bought rather than written
Version ordering is where this feature lives or dies, and its failure mode is
silent. `dpkg` ordering has epochs, and `~` sorts *before* the empty string, so
`3.0.2-0ubuntu1.15~rc1` precedes `3.0.2-0ubuntu1.15`. `rpmvercmp` has its own
segment rules and treats `~` and `^` differently again. A `strings.Compare` or
a semver parse orders `1.9` above `1.10` and reports a vulnerable fleet as
clean — a false negative, which nobody notices until it matters.
### Scanning one server
1. Load `server_packages`; resolve OS family and version to a `trivy-db`
ecosystem.
2. **Unsupported ecosystem → record `status: unsupported`, clear the flag,
write no findings.**
3. For each package: resolve source name, look up advisories, compare versions.
4. Upsert vulnerable results as `open`, preserving `first_seen`.
5. Any currently-`open` finding absent from this result set → `fixed`, stamp
`fixed_at`.
6. Any `accepted` finding past its `until` → back to `open`.
7. Clear `scan_pending`, stamp `scanned_at` and `db_version`.
Steps 5 and 6 must run in that order, so a finding that is both absent and
expired settles as `fixed` rather than reopening on a package that no longer
carries it.
Step 2 matters as much as any of the matching. Arch has no feed in `trivy-db`,
so an Arch host must report **unsupported**, never "0 findings". Reporting
clean when the truth is unknown is the same class of lie as a silently stale
database, and it is the reason `vulndb_meta.pulled_at` appears on screen rather
than only in a log.
### Severity
Resolved **vendor → NVD → unknown**, in that order, never invented.
This will surface as "why is this critical CVE marked low", and the answer is
that Debian and Red Hat routinely downgrade an NVD score because the vulnerable
code path is not reachable in their build. Their rating is the accurate one for
that package, and showing NVD's above it would manufacture work that does not
need doing.
---
## Findings lifecycle
`open | fixed | accepted`.
An accepted finding is suppressed from counts and alerts until its `until`
date, then reopens automatically. A reason is required.
Acceptance with a mandatory expiry, rather than permanent dismissal, is what
keeps the feature usable in both directions. Without any acceptance mechanism,
a kernel CVE awaiting a reboot window sits red indefinitely and trains people
to ignore the page. With permanent dismissal, accepted findings accumulate
silently and nobody revisits them — the dismissal list becomes where risk goes
to be forgotten, which is precisely what an auditor asks to see.
Retention: `settings.vuln_finding_retention_days`, a `*int` on the same pattern
as `workflow_log_retention_days` — nil means 90 days, 0 means forever. Only
`fixed` findings are swept, by a `StartVulnSweeper` inside the same
`RunAsLeader("housekeeping", …)` as the existing sweepers. `open` and
`accepted` findings are never swept at any setting.
---
## Alerting
Per-org rules over the existing `notification_channels`: severity threshold,
optional tag filter, target channels.
A rescan emits one message summarising what newly opened — "12 new critical
across 4 servers" — never one message per finding. See the leader section for
why the tick boundary makes this structural.
Modelling findings as a monitor type was rejected. It would reuse monitors'
state machine and channel wiring for free, but monitors are up/down for one
endpoint with retries and hourly rollups, none of which means anything for a
CVE; most fields would be disabled in the UI and the uptime graphs would be
polluted with a signal that is not uptime.
This adds one `notify` payload type and a `vuln_digest.html.tmpl` /
`vuln_digest.txt.tmpl` pair in `shared/mail`. Note that `shared/mail` templates
are parsed in `init()`, so a mistyped field is a boot-time panic — CLAUDE.md
describes a `render_test.go` guarding against exactly this, but **that file does
not exist**; the repository has no Go tests at all, and by instruction this
feature adds none. The template pair must therefore be verified by starting the
binary and sending one digest through a real channel.
---
## Entitlement
The feature name is `vuln_scanning`, and it crosses the two services the way
every other feature does:
- **admin** carries it as a per-instance entitlement toggle, so it can later be
priced as a catalogue `feature` component without a second migration;
- **the licence** snapshots it into `License.Features []string` at issue time;
- **the server** asks `lic.HasFeature("vuln_scanning")` and never switches on
tier, so changing what a tier includes needs no server release.
Off on Free.
**The gate is checked at `ReportPackages`, not at display.** Gating only the UI
would still pay every write cost, and storage is the expensive half.
The agent learns of it through the existing 30-second `SyncKeys` poll:
`SyncResponse` gains a `collect_packages` bool, and the hourly loop skips
collection entirely when it is false. So an ungated instance produces no
collection, no gRPC body, no document and no storage. `ReportPackages` still
re-checks the entitlement server-side and refuses — the agent flag is an
optimisation, the server check is the boundary.
Turning the feature off does not delete existing findings; they stop being
served and stop updating. Deletion is the instance-deletion path's job.
---
## REST API
```
GET /api/vulnerabilities # filter: severity, state, server, tags
GET /api/vulnerabilities/summary # severity counts + database freshness
POST /api/vulnerabilities/rescan # marks all scan_pending (owner|admin)
POST /api/vulnerabilities/:id/accept # reason + until (owner|admin)
DELETE /api/vulnerabilities/:id/accept # (owner|admin)
GET /api/servers/:id/vulnerabilities
GET /api/servers/:id/packages
GET /api/packages/search?name= # fleet-wide
GET,POST /api/vuln-rules · PUT,DELETE /api/vuln-rules/:id
```
Every mutating path writes an audit event, as all of them do. Acceptance is the
one decision people will be asked to justify, so `by`, `reason`, `until` and
`at` land in the audit record and not only on the document.
---
## UI
`/vulnerabilities` is a fleet board **grouped by CVE** — one row per CVE with
an affected-server count, expandable to the individual servers. The same CVE
across 40 servers is one decision, and a flat list of findings makes it look
like forty.
Server detail gains **Vulnerabilities** and **Packages** tabs. Alert rules go
on `/settings/notifications`, beside the channels they consume.
Remediation introduces no new mechanism: a finding carrying `fixed_in` renders
an **Apply updates** action calling the existing
`POST /api/servers/:id/apply-updates`, which is already `ApplyUpdatesCmd`. See
it, patch it, one place — and no second patching path to keep consistent with
the first.
Database freshness is shown wherever findings are, not tucked into settings. A
fleet scanning against a three-week-old database must say so rather than
quietly report all-clear.
---
## Verification
**No automated tests.** The repository has none today, and by explicit
instruction this feature adds none — no `*_test.go`, no frontend test files.
That is a deliberate decision by the repository owner, recorded here so the
absence reads as a choice rather than an omission.
It does change the risk profile, and the places it changes it are worth naming,
because each fails by producing a **wrong answer rather than a crash**:
- **Version comparison.** The backport case — installed `1:3.0.2-0ubuntu1.15`
against advisory fixed-in `1:3.0.2-0ubuntu1.15` resolving to *not
vulnerable* — plus tilde ordering (`1.0~rc1` < `1.0`), epoch dominance
(`1:1.0` > `2.0`) and `1.9` < `1.10`. Wrong here means a vulnerable fleet
reported clean.
- **Source-package fan-out.** One advisory against `openssl` must flag
`libssl3`, `openssl` and `libssl-dev`. Matching on binary name alone silently
finds one of three.
- **`first_seen` preservation.** An upsert that overwrites it makes every
finding look discovered today, and nothing surfaces that until someone reads
a report.
- **Fixed-before-reopen ordering.** A finding both absent from a scan and past
its acceptance expiry must settle `fixed`, not reopen.
The implementation plan carries a manual verification table for each, to be
walked before the relevant task is committed. They are the substitute for the
tests, not a formality.
---
## Failure modes
| Failure | Behaviour |
| ------- | --------- |
| Database pull fails | Keep the last good copy and serve stale. Record `last_error`, surface `pulled_at` age. **Never clear findings** — a network blip must not read as "all fixed" |
| Unsupported distribution | `status: unsupported`, not zero findings |
| Agent stops reporting | Findings persist and `collected_at` age is shown. No auto-expiry: a silent agent is not a patched server |
| `trivy-db` schema version bumps | The puller refuses an unknown schema rather than mis-parsing it |
| ghcr anonymous rate limit | Backoff; `VANTAGE_TRIVY_DB_REF` mirrors to a private registry |
| Leadership lost mid-scan | The context is cancelled and the scan returns; `scan_pending` is still set, so the next leader picks it up |
| Instance deleted | **`server_packages` and `vuln_findings` must be added to the control plane's instance-deletion collection list.** Easy to miss, and missing it orphans a tenant's package data indefinitely |
---
## Environment variables
| Name | Required | Notes |
| ---- | -------- | ----- |
| `VANTAGE_TRIVY_DB_REF` | no | default `ghcr.io/aquasecurity/trivy-db:2`. Point at a mirror for air-gapped installs or to avoid the anonymous ghcr rate limit |
| `VANTAGE_VULNDB_DISABLED` | no | disables the puller and the scheduler entirely. Findings already written are still served and still marked stale |
---
## Deliberately out of scope
- **Windows.** Separate source, collector and matcher; its own spec.
- **Container image scanning.** Sub-project C; needs the container registry.
- **Compliance baseline assertions.** Sub-project D; shares this findings UI
and nothing else.
- **Language-level dependency scanning** (npm, pip, Go modules). `trivy-db`
covers these ecosystems, but finding the manifests on a host is a different
collection problem from asking the package manager what is installed.
- **Automatic patching on a finding.** Remediation is one click, not zero. An
unattended upgrade triggered by a CVE feed is a fleet-wide change driven by a
third party's data, which is not a decision to take away from an operator.
@@ -1,441 +0,0 @@
# Workload registry
Date: 2026-08-06
Agents enumerate what each server actually runs — Docker containers, the
compose stacks grouping them, and systemd services — and report it to the
control plane. Containers and units can be started, stopped and restarted from
the UI, and a bounded snapshot of their logs can be read without opening a
console.
This is **sub-project B** of the four sketched in
`2026-08-06-package-inventory-and-cve-findings-design.md`:
| # | Sub-project | Depends on |
| - | ----------- | ---------- |
| A | Package inventory + CVE findings — its own spec | nothing |
| B | **Workload registry** — this spec | nothing |
| C | Container image scanning | A and B |
| D | Compliance profiles | shares A's findings UI only |
A and B are independent. C is the joiner and must not be designed before both
exist: it needs B's image list and A's findings model.
**Workload** is the domain word throughout: one container or one systemd unit.
It gives the collection, the commands and the page a single honest name rather
than saying "container or service" in every identifier.
Scope is **Linux only**, matching sub-project A and the existing position that
Windows agents are second-class by design. Docker runs on Windows; systemd does
not, and half a feature per platform is worse than a clear line.
---
## What this is for
The control plane can manage a fleet's keys, run workflows across it and watch
its endpoints, but it has no idea what any of those servers actually *runs*.
"Restart nginx on that box" means opening a console. "Which of these 80 servers
is still on the old image" is unanswerable.
---
## Reporting and refresh are one path
The agent reports on its own 60-second ticker through a `ReportWorkloads` RPC,
using the same hash short-circuit as the package report: it offers a SHA-256 of
the sorted workload list, and sends the body only when the server does not
already hold that hash. An unchanged list costs one small message, which on a
60-second cadence is the common case by a wide margin.
The on-demand refresh **does not return data**. `RefreshWorkloadsCmd` carries
no payload back; it makes the agent report immediately through the normal RPC,
and the UI refetches the stored document.
That is deliberate. A refresh that returned workloads inline would be a second
writer for the same collection, arriving by a different route, with its own
serialisation and its own opportunity to disagree with the periodic one. One
writer, one shape; the refresh is a nudge, not a channel.
Opening a server's Workloads tab dispatches a refresh, so what is on screen is
live rather than up to a minute stale. That matters because the page has a
Restart button on it: a stale list is not merely a wrong impression, it is a
wrong action aimed at a container that already died.
## What does answer back
Two operations genuinely return something:
| Command | Answers with |
| ------- | ------------ |
| `ControlWorkloadCmd{kind, id, action}` | the existing `CommandResult` — ok or error |
| `WorkloadLogsCmd{kind, id, tail}` | a new `WorkloadLogsResult{command_id, text, truncated}` |
Both ride the proven path: `commandDispatcher.send()` for request and ack, and
a `WorkloadResults` registry mirroring `StepResults.Await`/`Deliver` over the
bus. **`Await` must subscribe before the command is dispatched** — the pod
driving the request is usually not the pod holding the agent's stream, and a
fast agent otherwise answers into a channel nobody has joined. This is not a
new hazard; it is the one `stepresults.go` already documents.
```protobuf
rpc ReportWorkloads(ReportWorkloadsRequest) returns (ReportWorkloadsResponse);
message ReportWorkloadsRequest {
string server_id = 1;
string agent_token = 2;
string hash = 3;
bool docker_ok = 4;
string docker_error = 5;
bool systemd_ok = 6;
string systemd_error = 7;
repeated Workload workloads = 8; // empty on the offer call
}
message ReportWorkloadsResponse {
bool need_full = 1;
}
// ServerCommand gains three variants.
message RefreshWorkloadsCmd {}
message ControlWorkloadCmd {
string kind = 1; // "container" | "unit"
string id = 2;
string action = 3; // "start" | "stop" | "restart"
}
message WorkloadLogsCmd {
string kind = 1;
string id = 2;
int32 tail = 3;
}
// AgentMessage gains one variant.
message WorkloadLogsResult {
string command_id = 1;
string text = 2;
bool truncated = 3;
string error = 4;
}
```
The offer-then-send handshake is the package report's, unchanged: the agent
calls once with `workloads` empty, and resends with the body only if the
response sets `need_full`.
An agent whose stream no pod holds gets a 503 from the dispatcher, as
everything else does. Commands are not queued: a command whose owner died must
fail loudly rather than be delivered to nobody while the operator is told it
worked.
---
## Not gated by licence
Unlike CVE scanning, this reads as core fleet management rather than a premium
add-on, so v1 ships to every instance with no entitlement check.
If that changes it is a one-line `HasFeature` check at `ReportWorkloads`,
gating collection rather than display — the same placement and the same
reasoning as sub-project A, where gating the UI alone would still pay every
write cost.
---
## Data model
One new collection, `server_workloads`, one document per server, mirroring
`server_packages`.
```go
type ServerWorkloads struct {
ID bson.ObjectID `bson:"_id,omitempty" json:"-"`
InstanceID string `bson:"instance_id" json:"-"`
ServerID string `bson:"server_id" json:"server_id"`
Hash string `bson:"hash" json:"hash"`
Workloads []Workload `bson:"workloads" json:"workloads"`
CollectedAt time.Time `bson:"collected_at" json:"collected_at"`
DockerOK bool `bson:"docker_ok" json:"docker_ok"`
DockerError string `bson:"docker_error,omitempty" json:"docker_error,omitempty"`
SystemdOK bool `bson:"systemd_ok" json:"systemd_ok"`
SystemdError string `bson:"systemd_error,omitempty" json:"systemd_error,omitempty"`
}
type Workload struct {
Kind string `bson:"kind" json:"kind"` // "container" | "unit"
ID string `bson:"id" json:"id"` // container id, or unit name
Name string `bson:"name" json:"name"`
State string `bson:"state" json:"state"`
Health string `bson:"health,omitempty" json:"health,omitempty"`
Image string `bson:"image,omitempty" json:"image,omitempty"`
Stack string `bson:"stack,omitempty" json:"stack,omitempty"`
Ports []string `bson:"ports,omitempty" json:"ports,omitempty"`
Restarts int `bson:"restarts,omitempty" json:"restarts,omitempty"`
StartedAt time.Time `bson:"started_at,omitempty" json:"started_at,omitempty"`
Protected bool `bson:"protected" json:"protected"`
}
```
`State` is normalised across the two kinds: containers report `running`,
`exited`, `paused`, `restarting`, `created`; units report `active`, `inactive`,
`failed`, `activating`. They are deliberately **not** collapsed into a shared
vocabulary — a failed unit and an exited container mean different things, and
flattening them would lose the distinction the operator needs.
Indexes: `{instance_id, server_id}` unique, plus multikey
`{instance_id, "workloads.image"}` for the fleet-wide "which servers run image
X" query.
### Why the OK/Error pairs exist
A host with no Docker installed and a host where Docker is installed and
running nothing both produce an empty list. One should read "not in use here",
the other "nothing running", and only the second deserves any alarm.
The error strings separate a third case the booleans alone cannot: Docker
installed with the daemon down. "Not installed" and "installed but not
responding" are different problems with different fixes, and collapsing them
into one false boolean throws away the only thing that tells them apart.
### Why `Protected` is reported rather than derived
The agent already knows which unit and container it is. Sending that up lets
the UI render the action disabled with a reason instead of offering a button
whose refusal is already known.
The field is the courtesy; the agent's own check is the boundary. See the
control section.
### No history
A workload list is state, not a record. Nobody asks what containers ran last
Tuesday, and keeping it would grow a collection per server per minute in
exchange for a question nobody has.
---
## Collectors
### Docker: two commands, no English parsing
```
docker ps -aq
docker inspect --format '{{json .}}' <ids…>
```
Not `docker ps --format '{{json .}}'` alone. That reports health and uptime
inside a human `Status` string — `"Up 2 hours (healthy)"` — and anything built
on it is parsing English that is localised, reworded between releases, and
silently different for a paused or restarting container. `inspect` returns
`State.Health.Status`, `State.StartedAt` and `RestartCount` as typed fields.
Two execs instead of one, and no parser to be wrong.
`RestartCount` justifies the second call by itself: a container cycling is the
single thing this page most needs to show, and it is invisible in a list that
only ever says "Up".
Compose stacks come from the `com.docker.compose.project` label. **No YAML is
read from disk** — the label is what Docker itself treats as authoritative, and
a compose file on disk may not be what is actually running.
Docker absent, or a socket that cannot be reached, sets `DockerOK: false`. It
is not an error and produces no log line: most servers in a fleet built around
SSH key management will not have Docker, and treating the normal case as a
fault makes the feature look broken on the majority of the estate.
### systemd: filtered on purpose
```
systemctl list-units --type=service --state=running,failed --no-legend --plain --no-pager
systemctl list-unit-files --type=service --state=enabled --no-legend --plain --no-pager
```
Two calls because "running or failed" and "enabled but stopped" are different
questions, and an enabled unit that is not running is exactly the one worth
seeing.
Excluded by prefix: `systemd-`, `user@`, `session-`, `init.scope`. A typical
host carries 300+ units, the platform's own accounting for most of them.
Listing all of them buries the ten anyone cares about — the same failure mode
as an unfiltered vulnerability report, and the same fix.
Column output rather than `--output=json`: the JSON flag requires systemd 246+,
and this fleet includes older stable distributions. The column format has been
stable considerably longer than the JSON one has existed.
---
## Control actions
```
container: docker {start|stop|restart} <id>
unit: systemctl {start|stop|restart} <unit>
```
Owner or admin only. Every action writes an audit event naming the actor, the
server and the target.
### The protected set
Computed agent-side: `vantage-agent.service`, plus the container ID read from
`/proc/self/cgroup` should the agent ever be run inside a container.
The agent refuses those before doing anything. As with the console relay
hardcoding `127.0.0.1` agent-side, **the control plane may name a target, but
the agent decides what it will do to itself**. A server-side denylist alone
would be bypassed by the next dispatch path someone adds, and the failure is
unrecoverable from the UI: a server that stops its own agent goes offline, and
the way back is SSH or physical access — precisely what this feature exists to
avoid needing.
### Timeouts
`docker stop` waits on a container that may ignore SIGTERM. `systemctl stop`
on a unit with a long `TimeoutStopSec` blocks for exactly as long as that says.
Both run under a 90-second context, and a timeout returns a real error rather
than an ack implying success.
---
## Logs
```
container: docker logs --tail 500 --timestamps <id>
unit: journalctl -u <unit> -n 500 --no-pager --output=short-iso
```
Capped at **500 lines and 256KB, whichever binds first**, with `truncated` set
so the UI can say so. Two caps because 500 lines of a container emitting 4KB
JSON blobs is 2MB, and a line count alone does not stop it — the same reasoning
that gave workflow logs both a per-line and a per-run cap.
Live following is deliberately absent. The browser console already offers a
real terminal on the same server, where `docker logs -f` works properly with
its own scrollback and cancellation. Building a second streaming path — a
relay listener, proxy bus keys, a WebSocket upgrade and a cancellation story
for a follow nobody closed — to duplicate that would be a large amount of
machinery aimed at a capability already shipped. A bounded snapshot answers
"why did this restart", which is the question that sends people to the console
in the first place.
### Log reads are owner or admin only, and audited
Unlike workflow logs, these cannot be masked. A workflow's logs can be masked
because the run injected the secrets and therefore knows their values. A
container's stdout is arbitrary and may contain credentials nobody declared —
a connection string in a startup banner, a token in a stack trace.
So log reads sit behind the same role check as control actions and are audited.
A member who can see the fleet cannot read its logs. This is a deliberate
access decision, not an oversight, and it is why log reading is not simply
folded in with the read-only snapshot endpoints.
---
## REST API
```
GET /api/servers/:id/workloads # stored snapshot
POST /api/servers/:id/workloads/refresh # dispatch, then refetch
POST /api/servers/:id/workloads/:wid/action # {"action":"start|stop|restart"} (owner|admin)
GET /api/servers/:id/workloads/:wid/logs?tail= # (owner|admin)
GET /api/workloads?image=&stack=&state= # fleet-wide
```
`:wid` is a container ID or a unit name, URL-encoded. Unit names carry dots and
`@`, which are legal in a path segment but not worth relying on unencoded.
`tail` is clamped to the 500-line cap server-side; a client asking for more
gets 500, not an error.
---
## UI
Server detail gains a **Workloads** tab, ordered compose stacks first — grouped
under the stack name — then loose containers, then units.
That ordering is not cosmetic. A stack is one thing to an operator even when it
is six containers, and a flat list turns one decision into six rows. It is the
same argument that groups the vulnerabilities board by CVE rather than by
finding.
A `/workloads` fleet view answers "which servers run image X", which is the
reason the snapshot is stored at all rather than fetched on demand and
discarded.
Three rules that follow directly from the model:
- **Protected rows render their actions disabled, with the reason**, rather
than offering a button whose refusal is already known.
- **`DockerOK: false` reads "Docker not in use on this server"**, never an
empty list, and `DockerError` when present is shown as a distinct problem.
- State never reads by colour alone: every pill carries a distinct shape and a
text label, matching the existing monitor and severity pills.
---
## Verification
**No automated tests.** The repository has none today, and by explicit
instruction this feature adds none — no `*_test.go`, no frontend test files.
A deliberate decision by the repository owner, recorded so the absence reads as
a choice rather than an omission.
The behaviours that would otherwise have been tested are the ones that fail
quietly, and the implementation plan carries a manual check for each:
- **Parser output against real command output.** `docker inspect` must yield
`RestartCount`, health and the compose label as `Stack`; the `systemctl`
exclusion filter must drop `systemd-*` and `user@*` while keeping
`nginx.service`. Both are verified against a live host rather than a fixture.
- **Protected-set computation.** `vantage-agent.service` marked,
`nginx.service` not. Getting this wrong in the permissive direction lets a
server stop its own agent, which is unrecoverable from the UI.
- **Hash order-independence.** An ordering-sensitive hash resends the full list
every 60 seconds, which is invisible except as traffic.
- **Log capping in both directions.** 600 lines in → 500 out with `truncated`;
a 300KB blob of fewer than 500 lines → capped, `truncated`. The second is the
case a line-count-only implementation silently fails, and it fails by sending
megabytes rather than by erroring.
---
## Failure modes
| Failure | Behaviour |
| ------- | --------- |
| Docker not installed | `DockerOK: false`, no error, UI reads "not in use" |
| Docker installed, daemon down | `DockerOK: false` **plus** `DockerError` — different message, different fix |
| Agent offline | 503 from the existing dispatcher. No queueing: a command whose owner died must fail loudly |
| Action on a protected workload | Agent refuses; API answers 409 naming the reason |
| `stop` exceeds its timeout | Real error surfaced, never a hopeful ack. Snapshot refreshed afterwards |
| Container removed between snapshot and action | Docker's "No such container" surfaced and a refresh dispatched — this is what on-demand refresh is for |
| Log exceeds either cap | Truncated, flagged, and stated in the UI |
| Instance deleted | **`server_workloads` must be added to the control plane's instance-deletion collection list**, alongside sub-project A's two collections |
---
## Deliberately out of scope
- **Live log following.** The console already does it. See the logs section.
- **Creating, deleting or updating containers and units.** This is a control
and visibility surface, not a deployment tool — workflows already exist for
changing what a server runs, with snapshots, audit and rollback.
- **`docker exec` into a container.** The console reaches the host; exec from
the control plane is a second remote-execution path with its own audit and
authorisation story, and it belongs in its own spec if anywhere.
- **Kubernetes and containerd.** The Docker collector shells to the `docker`
CLI, so a node whose runtime is containerd or CRI-O reports nothing from it —
`DockerOK: false`, correctly, since Docker genuinely is not in use. Covering
those runtimes means a `crictl`/`nerdctl` collector, and talking to a
Kubernetes API server is a different subsystem again. Neither is v1.
- **Podman as a supported runtime.** Its `docker`-compatible CLI means an
aliased install will largely work, and that is a happy accident rather than a
claim: nothing here is tested against Podman and its `RestartCount` and
compose-label behaviour are not verified.
- **Windows.** No systemd, and a different container story.
- **Image vulnerability scanning.** Sub-project C, which needs this spec's
image list and sub-project A's findings model.
@@ -1,297 +0,0 @@
# API tokens and OpenAPI reference
Date: 2026-08-12
Status: approved, ready for implementation planning
## Problem
The only programmatic credential the control plane issues is the ESO secrets-read
bearer token, which reaches exactly one endpoint. Everything else requires a
browser session cookie. There is therefore no supported way to drive Vantage from
CI, a script, or infrastructure-as-code, and no machine-readable description of
the REST API for anyone who wants to try.
This spec covers two deliverables that ship together: scoped API tokens, and an
OpenAPI 3.1 document rendered as a live reference page. A Terraform provider is
the intended follow-on and is explicitly out of scope here — it depends on both
of these being settled, and it is a separate Go module with its own release
cycle.
## Goals
- A person can mint a scoped, optionally expiring token and use it against the
existing REST API with no new endpoints to learn.
- A leaked token is bounded by role, by scope, and by expiry policy.
- Offboarding a person removes their tokens as a side effect of removing them.
- The API has a machine-readable description that cannot silently drift from the
handlers it describes.
- The reference page works on an air-gapped self-hosted install.
## Non-goals
- Token editing. Role and scopes are immutable; rotation replaces amendment.
- OAuth device flow or any browser-based authorisation grant.
- Per-server or per-tag restrictions on a token.
- Instance-owned service tokens that outlive their creator.
- The Terraform provider.
- General API rate limiting beyond the per-token limit described below.
## Part 1 — API tokens
### Token format and storage
A token is `vt_` followed by 32 random bytes, hex encoded. It is displayed once,
at creation, and never again.
Only the SHA-256 hash is stored, in a unique index. This follows the precedent
already set by `servers.agent_token_hash` and the ESO read token. bcrypt is
deliberately not used: the value is full-entropy random rather than a
user-chosen password, so a fast hash is sufficient, and a per-token salt would
force a collection scan where an indexed lookup is wanted.
The first eight characters are stored in clear as `hint`, so the list can
identify a token without revealing it.
### Authentication path
`auth.Middleware()` gains a fallback. When there is no `km_session` cookie it
looks for `Authorization: Bearer vt_…`. Both paths end by placing a `*Session` in
the gin context, so every handler, `auth.RequireRole`, `RequireActiveLicense`,
`RequireFeature` and `actorFromCtx` continue to work unmodified.
```
Session{
UserID: token.UserID
InstanceID: token.InstanceID
Role: min(user.Role, token.Role) // owner > admin > member
Email: user.Email
TokenID: token.TokenID // "" for cookie sessions
Scopes: token.Scopes // nil for cookie sessions
}
```
The effective role is recomputed on every request rather than frozen at
creation. Demoting the user demotes the token with them. No caching is required
because the user document is already read to confirm the user still exists.
The existing host guard applies identically. A token carries an `instance_id`,
and a request arriving at a different instance's host is rejected exactly as a
mismatched cookie session is. The tenant boundary must not have a token-shaped
hole in it.
`last_used_at` is written best-effort and only when the stored value is more
than 60 seconds old, so it does not become a Mongo write per request.
Rejections:
| Condition | Status | Body |
| -------------------- | ------ | -------------------------------------- |
| No credential at all | 401 | `not authenticated` |
| Unknown token | 401 | `invalid token` |
| Expired token | 401 | `code: token_expired` |
| Owning user deleted | 401 | `invalid token` |
| Missing scope | 403 | names the required scope |
| Wrong instance host | 403 | `instance host mismatch` |
### Data model
New collection `api_tokens`, added to `services.ScopedCollections` so instance
purge reaches it.
```
instance_id string
token_id string
user_id string
name string 1-64 chars, unique per user
hint string first 8 chars of the plaintext
token_hash string sha256
role string owner|admin|member
scopes []string
expires_at *time.Time nil means never
created_at time.Time
last_used_at *time.Time
created_by_ip string
```
Indexes: unique on `token_hash`; compound on `(instance_id, user_id)`.
Deleting a user deletes their tokens as part of the same service call as
`DeleteInstanceUser`, so offboarding is one action rather than two.
### Expiry policy
Expiry is optional by default: a token may be created with no expiry at all.
Instance settings gain `api_token_max_days *int`, editable by owner and admin:
- `nil` — no cap; never-expire is allowed. This is the default, so an upgrade
changes nothing.
- `n > 0` — a new token must expire within `n` days, and a never-expire token is
refused.
Changing the setting does not retroactively invalidate existing tokens; it is a
policy on issuance. Tokens already outside the new cap are flagged in the UI so
that someone can rotate them deliberately, rather than discovering the change
when a pipeline breaks.
### Scopes
Eight resources, each with `:read` and `:write`. Write implies read on the same
resource.
```
servers keys secrets workflows
monitors vulns workloads settings
```
Scope enforcement is a single middleware, `RequireScopes()`, mounted once in the
`/api` stack. It derives the required resource from the matched gin route
pattern using a map, rather than from a per-route decorator: a route registered
without a decorator would otherwise be unguarded, and this repo already prefers
guards that come from where a route is mounted rather than from someone
remembering.
- Cookie sessions skip the check entirely.
- A token-authenticated request whose route pattern is absent from the map is
denied with 403. Fail closed.
- A startup check fails boot if any registered `/api` route pattern is missing
from the map, so the failure surfaces at deploy rather than at the first call.
Deliberate placements:
- `keys:read` covers `GET /keys/:id/private-key`. Reading a private key is
reading a key.
- `secrets:read` does not cover `GET /api/secrets/:group/values`. That endpoint
keeps its separate ESO bearer path and is unaffected by this work.
- `workloads:write` covers both container control actions and log reads, which
are already restricted to owner and admin.
- The token endpoints themselves map to the `settings` resource: `GET
/api/tokens` requires `settings:read`, and `POST` and `DELETE` require
`settings:write`. A token can therefore mint or revoke tokens only when
explicitly granted that scope, and never above its own role.
### Endpoints
```
GET /api/tokens list; a member sees their own, owner|admin see all
POST /api/tokens create; returns the plaintext once
DELETE /api/tokens/:id revoke; own always, owner|admin any
```
There is no `PUT`. Editing a token's role or scopes changes what a credential
already deployed in a CI system can do, with no record of what it could do
before. Rotation replaces amendment.
`POST` body: `name`, `role`, `scopes[]`, `expires_in_days` (omitted means never,
and is refused when `api_token_max_days` is set).
Refusals: 400 for an unknown scope, 409 for a duplicate name for that user, 403
for a role above the creator's own, 422 for an expiry beyond policy.
### Web UI
A new "API tokens" card in the Access group of `/settings`, alongside Members
and single sign-on. Not a new nav entry — `/settings/instance` was folded back
into `/settings` for precisely this reason, and the card lives in
`web/components/settings/` with the others, reusing the shared `Field` and
`inputClass`.
The card lists name, hint, role, scope chips, last used, and expiry with a
distinct state for expired and for over-policy. Revoke is per row and confirms.
Create opens a modal. The plaintext is shown once in a `--well` block with
copy-to-clipboard and an explicit line saying it will not be shown again.
Members see only their own rows. Owner and admin get an "All tokens" toggle.
`api_token_max_days` is a field on the same card, visible to owner and admin
only.
### Audit
New events:
- `token.created`
- `token.revoked`
- `token.expired_use` — a rejected expired token, which is how a forgotten CI
job becomes visible
- `settings.token_policy_updated`
The actor is the human's email throughout, so `actorFromCtx` needs no change.
Every existing audit event written during a token-authenticated request gains
`via: "token:<name>"` in its detail, so the log distinguishes a person clicking
from their credential acting.
### Rate limiting
Token-authenticated requests are limited per token in Redis at 600 per minute,
answering 429 with `Retry-After`. Cookie sessions are untouched. This is narrow
on purpose: it is not the general API rate-limiting project, only enough that a
runaway script cannot take an instance down.
## Part 2 — OpenAPI and the reference page
### Generation
`swaggo/swag` v2, pinned, emitting OpenAPI 3.1. v1 emits Swagger 2.0, which
Scalar renders poorly.
Handlers in `server/internal/api/*.go` gain annotation comments. Request and
response bodies that are currently anonymous inline structs become named
structs. This is real churn across roughly fifteen files and is the honest cost
of choosing generation over a hand-written document.
The generated `server/internal/api/docs/openapi.json` is committed and embedded
with `go:embed`, not generated during the image build: `server/Dockerfile`
produces a `scratch` runtime from a Go build stage, and adding codegen there
means putting the toolchain in the build image.
`server-deploy.yml` gains a check that regenerates the spec and runs
`git diff --exit-code`. An annotation edited without regenerating fails the
build. Without this check the annotations are worth less than a hand-written
document, because they would drift while appearing authoritative.
### Serving
```
GET /api/openapi.json the spec, session or token authenticated
GET /api/docs HTML page loading a vendored Scalar bundle
```
The Scalar standalone bundle is vendored under `server/internal/api/docs/`, with
its version recorded in a comment beside it and refreshed by hand. No CDN:
air-gapped self-hosted installs are supported, and a reference page that fails
closed on an offline site is a support ticket.
Because the page is served by the instance itself, "Try it" acts against the
reader's own API with their own session.
### Documented auth schemes
Three, kept distinct:
- `cookieAuth` — the `km_session` cookie.
- `bearerAuth` — a `vt_…` API token.
- The ESO secrets endpoint is marked as its own separate scheme, so nobody wires
a personal access token into External Secrets Operator.
## Documentation
- `docsite/docs/reference/api-tokens.md`: creating a token, the scope table,
curl examples, rotation, and the maximum-lifetime policy.
- `CLAUDE.md`: the three token routes under REST API, the `api_tokens`
collection, and a note that `openapi.json` is generated and CI-verified.
## Risks
- The anonymous-struct-to-named-struct conversion is the bulk of the work and
touches handler code this feature otherwise has no business in.
- The vendored Scalar bundle is a manual refresh that nobody will remember. The
version comment is the only mitigation.
- A scope map keyed on gin route patterns breaks if a route path is renamed. The
boot-time completeness check is what turns that into a startup failure rather
than a silent 403 in production.
## Follow-on work
A Terraform provider, as its own spec and plan, consuming the tokens and the
OpenAPI document produced here.
@@ -1,236 +0,0 @@
# Instance rename in Vantage HQ
**Date:** 2026-08-12
**Status:** approved, not yet implemented
## Problem
A cloud instance is named once, at creation, and never again. The name is
chosen in the first thirty seconds of a customer's relationship with the
product — before they have decided whether this is "Acme" or "Acme
Production" — and it is the name that becomes their DNS host, appears in every
sign-in link and heads every page of their control plane. Today the only way to
change it is to create a second instance and move, or to open a support ticket
that has no tooling behind it.
## What a rename is
One customer-initiated action on a **cloud** instance: a new name, from which a
new slug is derived, which moves the instance to a new DNS host.
Name and slug move together. The slug is re-derived through
`provision.BaseSlug`, so the rules that named the instance at creation are the
rules that rename it — the same reserved-label list, the same 340 character
bound, the same `Slugify` collapse of non-alphanumeric runs. There is no
separate slug field for the customer to edit, because two fields invite the
state where the name says one thing and the host says another, and that
divergence is exactly what a rename exists to fix.
A licence binds an instance **UUID**, not a slug. A rename therefore issues no
licence, calls Paddle not at all, and consumes no relink. This is the property
that makes the whole feature cheap, and it should be stated in any future change
that tempts someone to touch the licence from this path.
### What breaks, deliberately
- **The old host stops working.** The old slug is released the moment the rename
commits; another account may take it. Bookmarks, saved sign-in links and any
agent install one-liner that named the web host are stale. Agents themselves
are unaffected — they dial `GRPC_HOST`, which is not per-tenant.
- **The old host keeps working for up to 60 seconds.** `server/internal/auth/instancehost.go`
caches slug-to-instance lookups for 60s, and admin has no path to invalidate
another process's memory. The released slug can be claimed by another account
inside that window, so for up to a minute a replica still maps that host to the
previous tenant. No data is exposed — the host/session guard rejects a session
belonging to a different instance — but the new owner's users can briefly reach
the old tenant's instance on their own host, and see its login page rather than
theirs. Adding a cross-service invalidation channel for a 60-second window is
not worth the coupling.
- **The customer must sign in again.** `km_session` is set with no `Domain`
attribute, so it is host-only and does not follow the instance to its new
subdomain. The UI says so rather than letting the customer discover it.
## Scope
| | Customer (owner or admin) | Staff |
|---|---|---|
| Cloud instance | rename, 24h cooldown | rename, no cooldown |
| Self-hosted instance | refused, 400 | name only; there is no slug |
| Cloud placeholder | refused, 409 | refused, 409 |
Self-hosted is refused on the customer side for the same reason the member
endpoints refuse it: there is no control-plane row to write. The install is the
customer's, on their own database, and admin cannot reach it. Staff may still
correct the label on admin's own row, because that label is what staff search
by.
## Data flow
Two writes, in this order:
1. **Control plane `instances`**`{name, slug}`.
2. **Admin `admin_instances`**`{name, slug, renamed_at}`.
The control plane goes first because `instances.slug` carries the unique index,
and that index is what actually decides a race between two accounts reaching for
the same name. Deciding it anywhere else would be guessing.
If the second write fails, the first is rolled back best-effort — restoring the
previous name and slug — and the request answers 500. Leaving them divergent
would have HQ print a host that is not the host, which is worse than a failed
rename.
## Backend
### `shared/provision/instance.go`
```go
// ErrSlugTaken means the derived slug belongs to another instance.
var ErrSlugTaken = errors.New("slug taken")
// RenameInstance changes an instance's name and re-derives its slug.
func RenameInstance(ctx context.Context, db *mongo.Database, instanceID, name string) (*models.Instance, error)
```
It lives beside `CreateInstanceWithID` so slug derivation keeps one home, and it
behaves as that function's rules imply:
- `BaseSlug(name)` failures wrap `ErrNameRejected` — too short, too long,
reserved.
- The derived slug is compared against the instance's current one. If they are
equal, only the name is written; a cosmetic capitalisation change is not a
move, and must not fail on its own slug.
- **No `-2` suffix loop.** Creation appends a counter because the customer is
waiting on an instance and any free slug will do. A rename is a request for a
specific host, and silently landing the customer on `acme-2` is a worse answer
than refusing.
- A duplicate-key error on the update surfaces as `ErrSlugTaken`, exactly as the
create path treats it as "that slug is taken". The pre-check is a courtesy;
the index is the boundary.
### `admin/internal/cloudprov`
```go
func RenameInstance(ctx context.Context, instanceID, name string) (*sharedmodels.Instance, error)
```
A thin wrapper over `provision.RenameInstance` on `db.ControlDB()`. It writes
`instances` and nothing else, so admin's documented control-plane write boundary
`instances` and `users`, from `cloudprov` and `inject` only — is unchanged.
### `admin/internal/models`
`Instance` gains:
```go
// RenamedAt is when this instance last changed name, and backs the 24h
// customer cooldown. The cooldown is admin's policy, so it lives on admin's
// row rather than in the control plane, which has no opinion about how often
// a customer may move.
RenamedAt *time.Time `bson:"renamed_at,omitempty" json:"renamed_at,omitempty"`
```
A pointer because absent means "never renamed", and a zero `time.Time` would
read as 1 January year 1 — far enough in the past that the cooldown is inert,
but only by accident.
### `PUT /api/instances/:id/name` (customer)
Mounted in the `cust` group behind `auth.RequireAccountRole(owner, admin)`, and
resolving the instance through `ownedInstance` like every other instance route,
so another account's instance answers 404 rather than 403.
Body: `{"name": "..."}`, trimmed before use.
Refusals, in the order checked:
| Condition | Status | Body |
|---|---|---|
| `deployment != cloud` | 400 | `selfHostedRefusal`, the same constant and status the member endpoints already answer with |
| `placeholder` | 409 | instance is not provisioned yet |
| within 24h of `renamed_at` | 429 | includes the UTC time it unlocks |
| `provision.ErrNameRejected` | 422 | the wrapped reason, verbatim |
| `provision.ErrSlugTaken` | 409 | that name is already in use |
Success returns `{"instance_id", "name", "slug", "login_url"}` and writes an
audit entry `instance.renamed` with detail `<old-slug> -> <new-slug>`, so the
history of a host is answerable from the audit log alone.
`login_url` comes from the existing `loginURLFor(slug)`, which fills `{slug}`
into `APP_LOGIN_URL` — the same builder the licence emails already use, rather
than a second opinion about how a tenant host is spelled. It is empty when
`APP_LOGIN_URL` is unset, and the portal then falls back to the host string it
already composes from the slug in `InstanceRecord` and the instance page.
### `PUT /api/staff/instances/:id/name`
The same core, without the cooldown, actor recorded as the staff user. On a
self-hosted instance it updates `admin_instances.name` only and does not call
`cloudprov`.
## Frontend (`adminsite`)
### `lib/slug.ts`
A TypeScript mirror of `provision.Slugify` and the length/reserved checks, used
only to preview the resulting host while the customer types. It carries the same
warning as `web/lib/targets.ts`: it is a second implementation and must change in
the same commit as the Go one. The preview can disagree with the server — the
409 is the answer that counts.
### `components/RenamePanel.tsx`
An inline panel, not a modal — `adminsite` has no modal component, and the
codebase's idiom for a destructive-ish action with one input is `RelinkPanel`:
a control that expands in place inside a `Panel`.
Prefilled with the current name. Below the input, a live line reading
`acme-ltd.vantage.hostxtra.co.uk` as the customer types, and a note that they
will need to sign in again on the new host. Submit is disabled while the derived
slug is unchanged or invalid.
It lives in an "Address" panel on `app/(customer)/instances/[id]/page.tsx`,
rendered only when the instance is cloud and `account_role` is `owner` or
`admin`. The staff instance page mounts the same component against the staff
route.
`InstanceRecord` on the Overview page is not touched: it stays a summary, and
the rename is a decision that deserves the detail page.
### After a successful rename
Invalidate `["account"]`, collapse the panel, and let the page redraw with the new
name and host. The Console rail card shows the new host, with a note:
> This instance now lives at `acme-ltd.vantage.hostxtra.co.uk`. You will need to
> sign in again there.
**No automatic redirect.** Sending the browser to the new host lands the customer
on a login screen with no explanation, having just lost the HQ page they were
standing on. The link is right there; they click it when they are ready.
## Verification
The repository has no Go test suite, so verification is build plus manual
exercise, matching existing practice:
- `go build ./...` in `shared` and `admin`; `npm run build` in `adminsite`.
- Rename a cloud instance; confirm `instances` and `admin_instances` agree on
name and slug.
- The new host serves a login page; the old host stops resolving to the instance
within ~60 seconds.
- A second rename within 24 hours answers 429.
- A rename onto an occupied slug answers 409 and changes nothing.
- A rename attempt on a self-hosted instance from the customer portal answers
400, the same status and constant the member endpoints already answer with.
- The audit log carries `instance.renamed` with both slugs.
## Out of scope
- Slug aliases or redirects from the old host. The control plane resolves one
slug per instance, and an alias table is a second identity to keep correct for
the sake of stale bookmarks.
- Renaming from inside the control plane's own `/settings`. HQ owns instance
identity, the same way it owns licences and `hq`-sourced users; a second
writer would need the same collision handling and the same cooldown.
- Any change to the licence, subscription or Paddle line items.
@@ -0,0 +1,297 @@
# Windows agent parity: OS updates and workloads
Date: 2026-08-13
## Goal
Bring the Windows agent up to the Linux agent on two subsystems: OS update
check/apply, and the workload registry (collection, control, logs). Everything
else about the Windows agent stays as it is.
Out of scope, deliberately:
- **Package inventory and CVE findings.** `trivy-db` carries no Windows feed, so
a Windows finding needs a different source, a different matcher and a
different version comparison. That is its own project, and until it exists a
Windows host correctly reports `status: unsupported` rather than "0 findings".
- **SSH key management on Windows.** `administrators_authorized_keys` is a real
possibility but a separate decision.
- **winget.** Third-party app upgrades are a different question from OS
patching, and winget is absent on Server Core and older builds.
## Current state
The Windows agent registers, heartbeats, reports inventory, runs workflow steps
through PowerShell, relays console connections and self-updates via MSI. Four
gates stop it doing more:
| Gate | Location |
| --- | --- |
| `runtime.GOOS != "linux"` early return | `agent/internal/workloads/workloads.go`, `agent/internal/sync/workloads.go` (twice) |
| package-manager detection finds nothing | `agent/internal/updates/updates.go` (`detectPM`) |
| hard error | `agent/internal/packages/packages.go` (out of scope here) |
| `authorized_keys` write skipped | `agent/internal/sync/sync.go` (out of scope here) |
## Approach
The platform split moves into the agent, expressed as build tags following the
existing `inventory/collect_linux.go` / `collect_windows.go` /
`collect_other.go` precedent. The control plane stays OS-blind: `ReportWorkloads`,
`ControlWorkload`, `WorkloadLogs` and `ApplyUpdates` need no changes at all,
because a Windows service is reported as the same `unit` kind a systemd service
is.
Build tags rather than `runtime.GOOS` switches so PowerShell command strings do
not ship in the Linux binary, and so a platform left unimplemented is a compile
error rather than a silent no-op at runtime.
## Updates
### Layout
```
agent/internal/updates/
updates.go # PackageUpdate; CheckAvailable/ApplyAll declared once
updates_linux.go # existing detectPM, checkApt/DnfYum/Pacman/Zypper/Apk, ApplyAll
updates_windows.go # Windows Update COM, driven through PowerShell
updates_other.go # //go:build !linux && !windows — no-ops
```
`updates_other.go` carries the build constraint for the same reason
`inventory/collect_other.go` does: `_other` is not a GOOS suffix, so without the
constraint the file compiles everywhere and collides.
### Checking
One PowerShell invocation, `-NoProfile -NonInteractive`, emitting JSON:
```powershell
$searcher = (New-Object -ComObject Microsoft.Update.Session).CreateUpdateSearcher()
$result = $searcher.Search("IsInstalled=0 and Type='Software' and IsHidden=0")
```
The Windows Update COM API is used rather than the `PSWindowsUpdate` module
because it is present on every supported Windows, needs no PowerShell Gallery
install, and works unchanged against a WSUS server on an air-gapped fleet. The
agent runs as `LocalSystem`, which has the rights the API requires.
Each result maps to a `PackageUpdate`:
| Field | Value |
| --- | --- |
| `Name` | the update Title |
| `CurrentVersion` | empty |
| `NewVersion` | the KB article ID, e.g. `KB5034123` |
`CurrentVersion` is empty because a Windows update is not a version bump of a
named package, and inventing a current version would put a wrong string in front
of an operator. The KB ID goes in `NewVersion` because it is the identifier
people actually search for.
Timeout: 10 minutes. The first search after a boot contacts Microsoft Update and
is routinely slow.
### Applying
The same COM session: `CreateUpdateDownloader` then `CreateUpdateInstaller`,
over the updates returned by the search above, skipping any that require user
input. EULAs are accepted programmatically; an update whose EULA cannot be
accepted is skipped rather than failing the batch.
The operation fails when the installer's `ResultCode` is not 2 (succeeded) or 3
(succeeded with errors).
Timeout: 60 minutes. A patch-Tuesday cumulative genuinely takes that long, and
the Linux path's existing 5-minute cap is already tight.
**The agent never reboots the host.** A control plane silently restarting a
production server is unrecoverable from the UI, and the reboot is a decision a
person or a workflow makes. Instead the need for one is reported.
### Reboot required
A new field `reboot_required` on `InventoryReport`, added to
`proto/vantage/v1/vantage.proto` and to both hand-written `pb` copies
(`agent/internal/grpc/pb`, `server/internal/grpc/pb`) in the same commit.
It travels on the inventory report rather than the update report because it is a
host property like the kernel version, and it is set on the **static** snapshot
only — every 15 minutes rather than every 30 seconds. A host rebooted by hand
clears the flag in a quarter of an hour instead of showing it for up to a full
one, and the detection costs a PowerShell process on Windows, which is not
something to spawn twice a minute forever.
It is set in `agentsync.runInventory`, not inside the `inventory` package, so
`inventory` gains no dependency on `updates`.
Both platforms set it, since parity is free here:
- Linux: `/var/run/reboot-required` exists, or `dnf needs-restarting -r` exits
non-zero.
- Windows: the `Microsoft.Update.SystemInfo` COM object's `RebootRequired`
property, falling back to the pending-reboot registry keys
(`Component Based Servicing\RebootPending`,
`WindowsUpdate\Auto Update\RebootRequired`,
`Session Manager\PendingFileRenameOperations`).
`services.ReportInventory` stores it on `servers.inventory`.
## Workloads
### Layout
```
agent/internal/workloads/
workloads.go # Result, Collect, Hash — Collect calls collectUnits
docker.go # unchanged, shared: shells to the docker binary
systemd_linux.go # was systemd.go
services_windows.go # new: Win32_Service collection
control.go # shared validation; platform halves split out
control_linux.go # docker/systemctl, /proc/self/cgroup own-container check
control_windows.go # Start/Stop/Restart-Service, VantageAgent protection
logs.go # shared: capLog, MaxLogLines, MaxLogBytes
logs_linux.go # docker logs / journalctl
logs_windows.go # docker logs / Get-WinEvent
```
`Collect` loses its `runtime.GOOS != "linux"` return and calls
`collectUnits(ctx)`, which is the systemd collector on Linux and the service
collector on Windows. `runWorkloads` and `reportWorkloads` in
`agent/internal/sync/workloads.go` lose all three of their platform returns.
`docker.go` stays shared and ungated. It shells to the `docker` binary, which
behaves identically on Windows, so a Docker Desktop or Mirantis host reports its
containers with no new code. `DockerOK` / `DockerError` keep their existing
three-state meaning: not installed (the common case, not a fault), installed but
not responding, and running nothing.
### Collecting Windows services
`Get-CimInstance Win32_Service` converted to JSON — not `Get-Service`, which
exposes neither `PathName` nor `StartMode`, and the filter needs both.
A service is reported when its executable does **not** resolve under
`%SystemRoot%\System32`, and it is running, failed, or has `StartMode=Auto`
while stopped. This mirrors the systemd collector's intent: show what an
operator installed, and show what is meant to be up but is not.
Path parsing strips surrounding quotes and trailing arguments before the
`%SystemRoot%` comparison. `"C:\Program Files\X\x.exe" -service` is one path
with one argument, and splitting naively on whitespace misfiles a substantial
share of a real fleet.
Field mapping:
| Workload field | Source |
| --- | --- |
| `Kind` | `"unit"` |
| `ID` | `Name` (the service name) |
| `Name` | `DisplayName` |
| `State` | `running` / `stopped` / `failed`, from `State` plus `ExitCode` |
| `Health`, `Image`, `Stack`, `Ports`, `Restarts` | unset |
`Kind: "unit"` and the existing `systemd_ok` / `systemd_error` fields are reused
rather than a `service` kind and `services_ok` fields being added. That would
cost a proto change, both pb copies, the server model, the service layer and the
web client, and would teach every existing consumer a second kind — to describe
the same thing. The naming is corrected where it is read, in the UI, which knows
the server's OS.
`State` values match the ones the UI already colours, so no web change is needed
for the rows themselves.
### Protection
The protected set stays computed and enforced agent-side, as it is on Linux: the
control plane may name a target, but the agent decides what it will do to
itself.
On Windows the protected workload is the `VantageAgent` service — the NSSM
service name written by `installer/setup.ps1` — matched case-insensitively,
because Windows service names are. `detectOwnContainer` and its
`/proc/self/cgroup` read move to `control_linux.go`; the Windows build returns
no own-container ID.
`ErrProtected` still surfaces as HTTP 409 from the API, and the reported
`Protected` flag remains a courtesy that greys the button rather than the
boundary.
### Control
`Start-Service`, `Stop-Service -Force`, `Restart-Service -Force`, under the same
90-second `controlTimeout`, with the error text taken from PowerShell's stderr.
`sc.exe` is avoided because it returns before the operation completes, which
turns a timeout into a false success. `-Force` is required because
`Stop-Service` without it refuses when other services depend on the target, and
that refusal reads to an operator as a silent no-op.
### Logs
`Get-WinEvent` with a filter hashtable over the `System` and `Application` logs,
provider names matching the service name, its display name, and
`Service Control Manager`, newest first, capped by the requested tail.
Each event is formatted as `<ISO 8601 timestamp> <Level> <Message>`, which is
the same shape `journalctl --output=short-iso` produces, so the log dialog needs
no per-platform rendering.
Service Control Manager logs every service on the host under one provider, so
its events are filtered client-side to those naming the target service.
`capLog` is shared and unchanged: 500 lines and 256KB, whichever binds first,
trimmed from the front. There is still no follow mode.
An empty result returns an empty string and no error. A service that has logged
nothing is normal, and an error there would read as a broken feature.
## Server and web
The server changes in one place: `services.ReportInventory` persists
`reboot_required`.
The web changes in three, all keyed on the same `os_info` test
`MaintenanceTab.tsx` already uses (`server.os_info?.toLowerCase().includes("windows")`)
rather than on `os_type`. `os_type` is stored and serialised but unread by
`web/` today, and introducing a second Windows test in the same component is how
the two come to disagree. `WorkloadList` takes the result as a prop, since it
receives only a `serverId`:
1. `web/components/workloads/WorkloadList.tsx` — takes an `isWindows` prop from
the server detail page, and the systemd status lines become
platform-worded. On Windows the error line reads "Windows services could not
be read" and the "systemd is not in use on this server" line is not rendered
at all. The empty-state line drops "on Linux only". The Docker lines are
unchanged.
2. Server detail — a `Reboot required` pill beside the update count when the
flag is set, placed with the update panel because that is what caused it.
3. The Updates panel's Windows copy describes a list of KB articles rather than
package upgrades, since `current_version` is empty on that platform.
## Testing
The Windows collectors are, in substance, parsers of PowerShell output. Parsing
is separated from invocation and table-tested against captured real output. The
`agent` module has no tests at all today, so these are the first — they live
beside the parsers as ordinary `_test.go` files, run with `go test ./...` from
`agent/`, and need no new dependency:
- `Win32_Service` JSON, including a quoted path with arguments, a
`%SystemRoot%\System32` service that must be filtered out, a stopped
`StartMode=Auto` service that must be kept, and a failed service with a
non-zero `ExitCode`.
- Update searcher JSON, including an update with no KB ID.
- `Get-WinEvent` JSON, including a Service Control Manager event for another
service that must be filtered out.
- Pending-reboot detection from registry key presence.
`capLog` and `Hash` are unchanged and gain no tests.
The invocation halves are verified by hand on a Windows host: check, apply,
service start/stop/restart, a protected refusal on `VantageAgent`, and logs on
both a chatty service and a silent one.
`GOOS=windows go build ./...` and `GOOS=linux go build ./...` both belong in the
implementation plan as explicit steps — a build-tag split is exactly the change
that compiles on the machine you are sitting at and nowhere else. CI already
cross-builds the agent on release, so no workflow change is needed.
+5 -2
View File
@@ -55,8 +55,11 @@ It writes the config to `%ProgramData%\vantage\config.yaml`, installs the agent
as a Windows service and starts it.
:::info Windows servers do not get SSH key management
Windows agents register, report inventory and run workflow steps. Managing
`authorized_keys` is a Linux-only feature.
Windows agents register, heartbeat, report inventory, run workflow steps,
serve the browser console, check and apply OS updates, and report workloads
(services and containers). Managing `authorized_keys` is a Linux-only
feature, and so is package inventory and CVE scanning — the vulnerability
feeds this project uses carry no Windows data.
:::
## 3. Watch it come up
@@ -14,7 +14,7 @@ self-hosted; what differs is the term on offer, not what you get.
| | Free | Professional | Enterprise |
| --------------------- | --------- | ------------ | --------------------- |
| Servers (base) | 3 | 3 | 10 |
| Servers (base) | 3 | 5 | 10 |
| Monitors | 3 | unlimited | unlimited |
| Secret groups | 1 | unlimited | unlimited |
| Notification channels | 1 | unlimited | unlimited |
@@ -30,11 +30,11 @@ entitlement.
Three features are enabled per instance rather than bundled into a tier:
| Feature | What it enables |
| --------- | -------------------------------------------------------------------- |
| Browser console | The [browser console](../vantage/browser-console.md) |
| Single sign-on | [Sign-in through your identity provider](../vantage/settings.md#single-sign-on) |
| Vulnerability scanning| [Package vulnerability scanning](../vantage/vulnerabilities.md) |
| Feature | What it enables |
| ---------------------- | ------------------------------------------------------------------------------- |
| Browser console | The [browser console](../vantage/browser-console.md) |
| Single sign-on | [Sign-in through your identity provider](../vantage/settings.md#single-sign-on) |
| Vulnerability scanning | [Package vulnerability scanning](../vantage/vulnerabilities.md) |
No tier includes them by default; you enable them on the instances that need
them.
+11 -5
View File
@@ -89,10 +89,13 @@ metrics is normal rather than a fault.
### OS updates
Agents check for pending package updates hourly and report the count. From the
server page you can:
Agents check for pending package updates hourly and report the count the
machine's own package manager on Linux, the Windows Update COM API on Windows.
From the server page you can:
- **Apply updates** runs the machine's own package manager and reports back.
- **Apply updates** runs that check's install path and reports back. The agent
never reboots the machine; if one is owed, a **reboot required** badge
appears on the next inventory snapshot instead.
- **Update agent** upgrades the Vantage agent on that machine. See
[Agent updates](../operations/agent-updates.md).
@@ -107,8 +110,11 @@ Opens a browser SSH, RDP or VNC session. See [Browser console](./browser-console
## Windows servers
Windows agents register, run workflow steps and report inventory. They do not
manage `authorized_keys`.
Windows agents register, heartbeat, run workflow steps, report inventory,
check and apply OS updates, report workloads (services and containers), and
serve the browser console. They do not manage `authorized_keys`, and they are
not covered by package inventory or CVE scanning — the vulnerability feeds
this project uses carry no Windows data.
## Removing a server
+11 -10
View File
@@ -4,23 +4,24 @@ title: Workloads
sidebar_label: Workloads
---
A **workload** is one Docker container or one systemd service. Each Linux server
reports what it is running, and you can start, stop and restart those workloads,
and read their recent logs, without opening a console.
A **workload** is one Docker container or one service — a systemd unit on
Linux, a Windows service on Windows. Every server reports what it is running,
and you can start, stop and restart those workloads, and read their recent
logs, without opening a console.
Available on every instance. No licence feature is required.
## What gets reported
Linux servers only, reported every 60 seconds.
Every server, Linux and Windows, reported every 60 seconds.
- **Containers**: every container, running or not, with its image, published
ports, health, restart count and the compose stack it belongs to.
- **Services**: systemd units that are running, failed, or enabled but stopped.
The operating system's own units are hidden, since a typical host has hundreds
of them and they bury the ones you care about.
Windows servers report no workloads at all.
ports, health, restart count and the compose stack it belongs to. Requires
Docker (or Docker Desktop on Windows).
- **Services**: systemd units on Linux that are running, failed, or enabled but
stopped, and Windows services in the equivalent states. The operating
system's own units and platform services are hidden, since a typical host
has hundreds of them and they bury the ones you care about.
## Docker not in use is not an error
+3
View File
@@ -168,6 +168,9 @@ message InventoryReport {
uint64 swap_used = 7;
repeated PartitionReport partitions = 8;
string kernel = 9;
// Set on static snapshots only. The agent never reboots; it reports that one
// is owed and leaves the decision to a person or a workflow.
bool reboot_required = 10;
}
message InventoryReportResponse {
+10 -9
View File
@@ -113,15 +113,16 @@ type PartitionReport struct {
UsedBytes uint64 `json:"used_bytes"`
}
type InventoryReport struct {
ServerId string `json:"server_id"`
AgentToken string `json:"agent_token"`
IncludeStatic bool `json:"include_static"`
CPU *CPUReport `json:"cpu,omitempty"`
Memory *MemReport `json:"memory,omitempty"`
SwapTotal uint64 `json:"swap_total"`
SwapUsed uint64 `json:"swap_used"`
Partitions []PartitionReport `json:"partitions,omitempty"`
Kernel string `json:"kernel,omitempty"`
ServerId string `json:"server_id"`
AgentToken string `json:"agent_token"`
IncludeStatic bool `json:"include_static"`
CPU *CPUReport `json:"cpu,omitempty"`
Memory *MemReport `json:"memory,omitempty"`
SwapTotal uint64 `json:"swap_total"`
SwapUsed uint64 `json:"swap_used"`
Partitions []PartitionReport `json:"partitions,omitempty"`
Kernel string `json:"kernel,omitempty"`
RebootRequired bool `json:"reboot_required,omitempty"`
}
type InventoryReportResponse struct{}
+1
View File
@@ -39,6 +39,7 @@ type Inventory struct {
SwapUsedBytes uint64 `bson:"swap_used_bytes" json:"swap_used_bytes"`
Partitions []Partition `bson:"partitions,omitempty" json:"partitions,omitempty"`
Kernel string `bson:"kernel,omitempty" json:"kernel,omitempty"`
RebootRequired bool `bson:"reboot_required,omitempty" json:"reboot_required,omitempty"`
MetricsAt *time.Time `bson:"metrics_at,omitempty" json:"metrics_at,omitempty"`
StaticAt *time.Time `bson:"static_at,omitempty" json:"static_at,omitempty"`
}
+1
View File
@@ -28,6 +28,7 @@ func StoreInventory(serverID string, r *pb.InventoryReport) error {
set["inventory.static_at"] = now
set["inventory.swap_total_bytes"] = r.SwapTotal
set["inventory.kernel"] = r.Kernel
set["inventory.reboot_required"] = r.RebootRequired
if r.CPU != nil {
set["inventory.cpu.model"] = r.CPU.Model
set["inventory.cpu.cores"] = r.CPU.Cores
+3 -1
View File
@@ -301,7 +301,9 @@ export default function ServerDetailPage() {
<div ref={panelsRef} className="p-4 sm:p-6 lg:p-8">
<div role="tabpanel" id={`server-panel-${activeTab}`} aria-labelledby={`server-tab-${activeTab}`}>
{activeTab === "overview" && <OverviewTab server={server} attention={attention} onGoTo={selectTab} />}
{activeTab === "workloads" && <WorkloadList serverId={server.server_id} canControl={isAdmin} />}
{activeTab === "workloads" && (
<WorkloadList serverId={server.server_id} canControl={isAdmin} isWindows={!!server.os_info?.toLowerCase().includes("windows")} />
)}
{activeTab === "security" && <ServerVulnerabilities serverId={server.server_id} />}
{activeTab === "access" && <AccessTab server={server} onGenerateKey={() => setShowGenerateModal(true)} />}
{activeTab === "maintenance" && (
+13 -6
View File
@@ -47,11 +47,18 @@ export function MaintenanceTab({
<Card padding={false}>
<div className="flex flex-wrap items-center justify-between gap-3 border-b border-border px-6 py-4">
<h2 className="text-lg font-semibold text-text-primary">OS updates</h2>
{updates.length > 0 ? <Badge variant="warning">{updates.length} pending</Badge> : <Badge variant="success">up to date</Badge>}
<div className="flex flex-wrap items-center gap-2">
{updates.length > 0 ? <Badge variant="warning">{updates.length} pending</Badge> : <Badge variant="success">up to date</Badge>}
{/* Sits with the updates panel because that is what
caused it. The agent never reboots a host itself. */}
{server.inventory?.reboot_required && <Badge variant="warning">reboot required</Badge>}
</div>
</div>
{updates.length === 0 ? (
<p className="px-6 py-10 text-center text-sm text-text-secondary">No pending package updates. The agent checks hourly.</p>
<p className="px-6 py-10 text-center text-sm text-text-secondary">
{isWindows ? "No pending Windows updates. The agent checks hourly." : "No pending package updates. The agent checks hourly."}
</p>
) : (
<>
{/* Capped so a host with 400 pending packages does not make the
@@ -60,21 +67,21 @@ export function MaintenanceTab({
<Table>
<Thead>
<Tr>
<Th>Package</Th>
<Th>{isWindows ? "Update" : "Package"}</Th>
<Th>Current</Th>
<Th>Available</Th>
<Th>{isWindows ? "KB" : "Available"}</Th>
</Tr>
</Thead>
<Tbody>
{updates.map((u) => (
<Tr key={u.name}>
<Td label="Package">
<Td label={isWindows ? "Update" : "Package"}>
<span className="font-mono text-sm font-medium">{u.name}</span>
</Td>
<Td label="Current">
<span className="font-mono text-xs text-text-secondary">{u.current_version || "n/a"}</span>
</Td>
<Td label="Available">
<Td label={isWindows ? "KB" : "Available"}>
<span className="font-mono text-xs text-success">{u.new_version}</span>
</Td>
</Tr>
+13 -4
View File
@@ -45,7 +45,7 @@ const ACTION_PAST: Record<WorkloadAction, string> = {
restart: "Restarted",
};
export function WorkloadList({ serverId, canControl }: { serverId: string; canControl: boolean }) {
export function WorkloadList({ serverId, canControl, isWindows }: { serverId: string; canControl: boolean; isWindows: boolean }) {
const qc = useQueryClient();
const toast = useToast();
const [logTarget, setLogTarget] = useState<Workload | null>(null);
@@ -115,7 +115,7 @@ export function WorkloadList({ serverId, canControl }: { serverId: string; canCo
{snapshot.isLoading ? (
<p className="text-sm text-text-secondary">Loading</p>
) : !data ? (
<p className="text-sm text-text-secondary">Nothing reported yet. Agents report every 60 seconds, on Linux only.</p>
<p className="text-sm text-text-secondary">Nothing reported yet. Agents report every 60 seconds.</p>
) : (
<div className="space-y-2 text-sm">
{/* Docker absent is the common case on a fleet built
@@ -128,9 +128,18 @@ export function WorkloadList({ serverId, canControl }: { serverId: string; canCo
<p className="text-text-secondary">Docker is not in use on this server.</p>
) : null}
{/* One wire field, two honest words for it: the agent
reports Windows services under the same `unit` kind
systemd units use, and only the UI knows which host
this is. On Windows there is no "not in use" case
every Windows host has a service controller so a
failure is the only thing worth saying. */}
{data.systemd_error ? (
<p className="text-warning">systemd could not be read: {data.systemd_error}</p>
) : !data.systemd_ok ? (
<p className="text-warning">
{isWindows ? "Windows services could not be read: " : "systemd could not be read: "}
{data.systemd_error}
</p>
) : !data.systemd_ok && !isWindows ? (
<p className="text-text-secondary">systemd is not in use on this server.</p>
) : null}
</div>
+1
View File
@@ -14,6 +14,7 @@ export interface Inventory {
swap_used_bytes: number;
partitions?: { device: string; mountpoint: string; fstype?: string; total_bytes: number; used_bytes: number }[];
kernel?: string;
reboot_required?: boolean;
metrics_at?: string;
static_at?: string;
}