diff --git a/docs/superpowers/plans/2026-07-20-fleet-inventory.md b/docs/superpowers/plans/2026-07-20-fleet-inventory.md deleted file mode 100644 index 1788583..0000000 --- a/docs/superpowers/plans/2026-07-20-fleet-inventory.md +++ /dev/null @@ -1,866 +0,0 @@ -# Fleet Inventory Implementation Plan - -> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. - -**Goal:** Agents collect CPU/RAM/swap/disk/partition inventory and report it to the server via a new `ReportInventory` RPC; the server stores the latest snapshot per server and the UI displays it. - -**Architecture:** New unary gRPC `ReportInventory` (mirrors existing `ReportUpdates`). Agent runs a 30s metrics ticker (CPU/RAM/swap usage) and, every 15 min, a full static collection (disks, partitions, CPU model, kernel). Server upserts an embedded `inventory` sub-doc on the `servers` document with merge rules that preserve static fields between slow ticks. - -**Tech Stack:** Go (gin, mongo-driver v2, hand-written JSON-codec gRPC), `/proc` readers, Next.js 16 + react-query + Tailwind. - -## Global Constraints - -- **No tests this iteration.** Verify with `go build ./...`, `go vet ./...`, `npm run build`. -- gRPC uses a JSON codec: edit **both** `server/internal/grpc/pb/vantage.pb.go` and `agent/internal/grpc/pb/vantage.pb.go` identically, plus `proto/vantage/v1/vantage.proto` as documentation. No codegen. Mirror the existing `ReportUpdates` RPC wiring exactly (service interface, `_Vantage_*_Handler`, client method, `Vantage_ServiceDesc`). -- Mongo: `db.Col("servers")`, `context.WithTimeout`. Follow `server/internal/services/servers.go`. -- Agent already runs as root; `/proc` is readable. Linux is primary; Windows collectors may return empty. -- Module path `github.com/mrhid6/vantage`. -- Do not add heavy dependencies; implement `/proc` parsing directly. - ---- - -## Task 1: Inventory model + gRPC messages - -**Files:** -- Modify: `server/internal/models/server.go` -- Modify: `proto/vantage/v1/vantage.proto` -- Modify: `server/internal/grpc/pb/vantage.pb.go` -- Modify: `agent/internal/grpc/pb/vantage.pb.go` - -**Interfaces:** -- Produces: `models.Inventory` (+ `CPUInfo`, `MemInfo`, `Partition`) and `Server.Inventory *Inventory`. pb structs `InventoryReport`, `CPUReport`, `MemReport`, `PartitionReport`, `InventoryReportResponse`. Service method `ReportInventory` on both client and server interfaces. - -- [ ] **Step 1: Add model structs** - -In `server/internal/models/server.go` add (keep the existing `import "time"`): - -```go -type CPUInfo struct { - Model string `bson:"model,omitempty" json:"model,omitempty"` - Cores int `bson:"cores,omitempty" json:"cores,omitempty"` - UsagePct float64 `bson:"usage_pct" json:"usage_pct"` - Load1 float64 `bson:"load1,omitempty" json:"load1,omitempty"` -} - -type MemInfo struct { - TotalBytes uint64 `bson:"total_bytes" json:"total_bytes"` - UsedBytes uint64 `bson:"used_bytes" json:"used_bytes"` -} - -type Partition struct { - Device string `bson:"device" json:"device"` - Mountpoint string `bson:"mountpoint" json:"mountpoint"` - Fstype string `bson:"fstype,omitempty" json:"fstype,omitempty"` - TotalBytes uint64 `bson:"total_bytes" json:"total_bytes"` - UsedBytes uint64 `bson:"used_bytes" json:"used_bytes"` -} - -type Inventory struct { - CPU CPUInfo `bson:"cpu" json:"cpu"` - Memory MemInfo `bson:"memory" json:"memory"` - SwapTotalBytes uint64 `bson:"swap_total_bytes" json:"swap_total_bytes"` - SwapUsedBytes uint64 `bson:"swap_used_bytes" json:"swap_used_bytes"` - Partitions []Partition `bson:"partitions,omitempty" json:"partitions,omitempty"` - Kernel string `bson:"kernel,omitempty" json:"kernel,omitempty"` - MetricsAt *time.Time `bson:"metrics_at,omitempty" json:"metrics_at,omitempty"` - StaticAt *time.Time `bson:"static_at,omitempty" json:"static_at,omitempty"` -} -``` - -Add to the `Server` struct: `Inventory *Inventory \`bson:"inventory,omitempty" json:"inventory,omitempty"\``. - -- [ ] **Step 2: Document RPC in proto** - -In `proto/vantage/v1/vantage.proto`, add to the service: `rpc ReportInventory(InventoryReport) returns (InventoryReportResponse);` and the messages `InventoryReport`, `CPUReport`, `MemReport`, `PartitionReport`, `InventoryReportResponse` per spec §4. - -- [ ] **Step 3: Add pb structs + RPC wiring (server pb)** - -In `server/internal/grpc/pb/vantage.pb.go` add the message structs: - -```go -type CPUReport struct { - Model string `json:"model,omitempty"` - Cores int `json:"cores,omitempty"` - UsagePct float64 `json:"usage_pct"` - Load1 float64 `json:"load1,omitempty"` -} -type MemReport struct { - TotalBytes uint64 `json:"total_bytes"` - UsedBytes uint64 `json:"used_bytes"` -} -type PartitionReport struct { - Device string `json:"device"` - Mountpoint string `json:"mountpoint"` - Fstype string `json:"fstype,omitempty"` - TotalBytes uint64 `json:"total_bytes"` - UsedBytes uint64 `json:"used_bytes"` -} -type InventoryReport struct { - ServerId string `json:"server_id"` - AgentToken string `json:"agent_token"` - IncludeStatic bool `json:"include_static"` - CPU *CPUReport `json:"cpu,omitempty"` - Memory *MemReport `json:"memory,omitempty"` - SwapTotal uint64 `json:"swap_total"` - SwapUsed uint64 `json:"swap_used"` - Partitions []PartitionReport `json:"partitions,omitempty"` - Kernel string `json:"kernel,omitempty"` -} -type InventoryReportResponse struct{} -``` - -Then mirror the `ReportUpdates` RPC plumbing for `ReportInventory`. Locate every `ReportUpdates` reference in this file and add the parallel `ReportInventory`: -- `VantageServer` interface: add `ReportInventory(context.Context, *InventoryReport) (*InventoryReportResponse, error)`. -- `UnimplementedVantageServer`: add the stub returning `Unimplemented`. -- `VantageClient` interface + `keyManagerClient`: add the client method `Invoke`-ing `/vantage.v1.Vantage/ReportInventory`. -- `Vantage_ServiceDesc.Methods`: add `{MethodName: "ReportInventory", Handler: _Vantage_ReportInventory_Handler}`. -- Add `_Vantage_ReportInventory_Handler` copied from `_Vantage_ReportUpdates_Handler` with types swapped. - -- [ ] **Step 4: Mirror pb structs + wiring (agent pb)** - -Apply the identical additions to `agent/internal/grpc/pb/vantage.pb.go`. - -- [ ] **Step 5: Verify build** - -Run: `cd server && go build ./... && cd ../agent && go build ./...` -Expected: both succeed. - -- [ ] **Step 6: Commit** - -```bash -git add server/internal/models/server.go proto/vantage/v1/vantage.proto server/internal/grpc/pb/vantage.pb.go agent/internal/grpc/pb/vantage.pb.go -git commit -m "feat(proto): add ReportInventory RPC and inventory model" -``` - ---- - -## Task 2: Server handler + store service - -**Files:** -- Create: `server/internal/services/inventory.go` -- Modify: `server/internal/grpc/server.go` - -**Interfaces:** -- Consumes: `pb.InventoryReport` (T1), `db.Col("servers")`. -- Produces: `services.StoreInventory(serverID string, r *pb.InventoryReport) error`; gRPC method `(*vantageServer).ReportInventory`. - -- [ ] **Step 1: Write the store service** - -```go -package services - -import ( - "context" - "time" - - "github.com/mrhid6/vantage/server/internal/db" - "github.com/mrhid6/vantage/server/internal/grpc/pb" - "go.mongodb.org/mongo-driver/v2/bson" -) - -// StoreInventory upserts the latest inventory snapshot onto the server document. -// Metrics fields update every call; static fields only when r.IncludeStatic. -func StoreInventory(serverID string, r *pb.InventoryReport) error { - ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) - defer cancel() - - now := time.Now() - set := bson.M{"inventory.metrics_at": now} - if r.CPU != nil { - set["inventory.cpu.usage_pct"] = r.CPU.UsagePct - set["inventory.cpu.load1"] = r.CPU.Load1 - } - if r.Memory != nil { - set["inventory.memory.used_bytes"] = r.Memory.UsedBytes - } - set["inventory.swap_used_bytes"] = r.SwapUsed - - if r.IncludeStatic { - set["inventory.static_at"] = now - set["inventory.swap_total_bytes"] = r.SwapTotal - set["inventory.kernel"] = r.Kernel - if r.CPU != nil { - set["inventory.cpu.model"] = r.CPU.Model - set["inventory.cpu.cores"] = r.CPU.Cores - } - if r.Memory != nil { - set["inventory.memory.total_bytes"] = r.Memory.TotalBytes - } - parts := make([]bson.M, 0, len(r.Partitions)) - for _, p := range r.Partitions { - parts = append(parts, bson.M{ - "device": p.Device, "mountpoint": p.Mountpoint, "fstype": p.Fstype, - "total_bytes": p.TotalBytes, "used_bytes": p.UsedBytes, - }) - } - set["inventory.partitions"] = parts - } - - _, err := db.Col("servers").UpdateOne(ctx, bson.M{"server_id": serverID}, bson.M{"$set": set}) - return err -} -``` - -- [ ] **Step 2: Add the gRPC handler** - -In `server/internal/grpc/server.go`, add (mirroring the existing `ReportUpdates` handler that validates the agent token): - -```go -func (s *vantageServer) ReportInventory(ctx context.Context, req *pb.InventoryReport) (*pb.InventoryReportResponse, error) { - srv, err := services.ValidateAgentToken(req.ServerId, req.AgentToken) - if err != nil { - return nil, status.Errorf(codes.Unauthenticated, "invalid agent token") - } - if err := services.StoreInventory(srv.ServerID, req); err != nil { - log.Printf("store inventory for %s: %v", srv.ServerID, err) - } - return &pb.InventoryReportResponse{}, nil -} -``` - -Confirm `status`, `codes`, `log` are already imported in the file (they are, used by other handlers). - -- [ ] **Step 3: Verify build** - -Run: `cd server && go build ./... && go vet ./...` -Expected: success. - -- [ ] **Step 4: Commit** - -```bash -git add server/internal/services/inventory.go server/internal/grpc/server.go -git commit -m "feat(server): store inventory and handle ReportInventory RPC" -``` - ---- - -## Task 3: Agent collectors - -**Files:** -- Create: `agent/internal/inventory/collect_linux.go` -- Create: `agent/internal/inventory/collect_other.go` -- Create: `agent/internal/inventory/inventory.go` - -**Interfaces:** -- Produces: `inventory.Collect(includeStatic bool) *pb.InventoryReport`. - -- [ ] **Step 1: Common entry (`inventory.go`)** - -```go -package inventory - -import "github.com/mrhid6/vantage/agent/internal/grpc/pb" - -// Collect gathers metrics always and static hardware info when includeStatic. -// Platform specifics are provided by collect_linux.go / collect_other.go. -func Collect(includeStatic bool) *pb.InventoryReport { - r := &pb.InventoryReport{IncludeStatic: includeStatic, CPU: &pb.CPUReport{}, Memory: &pb.MemReport{}} - collect(r, includeStatic) - return r -} -``` - -- [ ] **Step 2: Linux collector (`collect_linux.go`)** - -Build-tagged `//go:build linux`. Implement `collect(r *pb.InventoryReport, includeStatic bool)`: -- CPU usage: read `/proc/stat` first line twice ~100ms apart, compute `1 - idleDelta/totalDelta` × 100 → `r.CPU.UsagePct`. -- Load: first field of `/proc/loadavg` → `r.CPU.Load1`. -- Mem/swap: parse `/proc/meminfo` (`MemTotal`, `MemAvailable`, `SwapTotal`, `SwapFree`; used = total − available; swap used = swaptotal − swapfree) → `r.Memory.*`, `r.SwapUsed`, and on static `r.SwapTotal`. -- Static only: `/proc/cpuinfo` (`model name`, count `processor` lines) → `r.CPU.Model/Cores`; `/proc/meminfo MemTotal` → `r.Memory.TotalBytes`; kernel via `syscall.Uname` or read `/proc/sys/kernel/osrelease` → `r.Kernel`; partitions from `/proc/mounts` filtered to fstypes in {ext4,xfs,btrfs,zfs,vfat,ntfs} then `syscall.Statfs` for total/used → `r.Partitions`. - -```go -//go:build linux - -package inventory - -import ( - "bufio" - "os" - "strconv" - "strings" - "syscall" - "time" - - "github.com/mrhid6/vantage/agent/internal/grpc/pb" -) - -func collect(r *pb.InventoryReport, includeStatic bool) { - r.CPU.UsagePct = cpuUsage() - r.CPU.Load1 = load1() - memTotal, memAvail, swapTotal, swapFree := meminfo() - if memTotal > memAvail { - r.Memory.UsedBytes = memTotal - memAvail - } - if swapTotal > swapFree { - r.SwapUsed = swapTotal - swapFree - } - if includeStatic { - r.Memory.TotalBytes = memTotal - r.SwapTotal = swapTotal - r.CPU.Model, r.CPU.Cores = cpuStatic() - r.Kernel = kernel() - r.Partitions = partitions() - } -} - -func readProc(path string) string { b, _ := os.ReadFile(path); return string(b) } - -func cpuSample() (idle, total uint64) { - f, err := os.Open("/proc/stat") - if err != nil { - return - } - defer f.Close() - sc := bufio.NewScanner(f) - if sc.Scan() { - fields := strings.Fields(sc.Text()) // cpu user nice system idle iowait ... - for i, v := range fields[1:] { - n, _ := strconv.ParseUint(v, 10, 64) - total += n - if i == 3 { // idle - idle = n - } - } - } - return -} - -func cpuUsage() float64 { - i1, t1 := cpuSample() - time.Sleep(100 * time.Millisecond) - i2, t2 := cpuSample() - dt := float64(t2 - t1) - if dt <= 0 { - return 0 - } - return (1 - float64(i2-i1)/dt) * 100 -} - -func load1() float64 { - fields := strings.Fields(readProc("/proc/loadavg")) - if len(fields) > 0 { - v, _ := strconv.ParseFloat(fields[0], 64) - return v - } - return 0 -} - -func meminfo() (total, avail, swapTotal, swapFree uint64) { - f, err := os.Open("/proc/meminfo") - if err != nil { - return - } - defer f.Close() - sc := bufio.NewScanner(f) - for sc.Scan() { - fields := strings.Fields(sc.Text()) - if len(fields) < 2 { - continue - } - kb, _ := strconv.ParseUint(fields[1], 10, 64) - b := kb * 1024 - switch strings.TrimSuffix(fields[0], ":") { - case "MemTotal": - total = b - case "MemAvailable": - avail = b - case "SwapTotal": - swapTotal = b - case "SwapFree": - swapFree = b - } - } - return -} - -func cpuStatic() (model string, cores int) { - f, err := os.Open("/proc/cpuinfo") - if err != nil { - return - } - defer f.Close() - sc := bufio.NewScanner(f) - for sc.Scan() { - line := sc.Text() - if strings.HasPrefix(line, "processor") { - cores++ - } else if strings.HasPrefix(line, "model name") && model == "" { - if i := strings.Index(line, ":"); i >= 0 { - model = strings.TrimSpace(line[i+1:]) - } - } - } - return -} - -func kernel() string { - return strings.TrimSpace(readProc("/proc/sys/kernel/osrelease")) -} - -func partitions() []pb.PartitionReport { - allowed := map[string]bool{"ext4": true, "xfs": true, "btrfs": true, "zfs": true, "vfat": true, "ntfs": true, "ext3": true} - f, err := os.Open("/proc/mounts") - if err != nil { - return nil - } - defer f.Close() - var out []pb.PartitionReport - seen := map[string]bool{} - sc := bufio.NewScanner(f) - for sc.Scan() { - fields := strings.Fields(sc.Text()) - if len(fields) < 3 || !allowed[fields[2]] || seen[fields[1]] { - continue - } - seen[fields[1]] = true - var st syscall.Statfs_t - if syscall.Statfs(fields[1], &st) != nil { - continue - } - total := st.Blocks * uint64(st.Bsize) - free := st.Bavail * uint64(st.Bsize) - out = append(out, pb.PartitionReport{ - Device: fields[0], Mountpoint: fields[1], Fstype: fields[2], - TotalBytes: total, UsedBytes: total - free, - }) - } - return out -} -``` - -- [ ] **Step 3: Non-linux stub (`collect_other.go`)** - -```go -//go:build !linux - -package inventory - -import "github.com/mrhid6/vantage/agent/internal/grpc/pb" - -// collect is a no-op best-effort stub on non-Linux platforms. -func collect(r *pb.InventoryReport, includeStatic bool) {} -``` - -- [ ] **Step 4: Verify build** - -Run: `cd agent && go build ./... && go vet ./...` -Expected: success (build both native and, if convenient, `GOOS=windows go build ./...`). - -- [ ] **Step 5: Commit** - -```bash -git add agent/internal/inventory/ -git commit -m "feat(agent): /proc-based inventory collectors" -``` - ---- - -## Task 4: Agent client method + scheduler - -**Files:** -- Modify: `agent/internal/grpc/client.go` -- Modify: the agent main loop (`agent/cmd/main.go` or `agent/internal/sync/sync.go` — wherever the poll loop/tickers live). - -**Interfaces:** -- Consumes: `inventory.Collect` (T3), pb (T1). -- Produces: `(*Client).ReportInventory(report *pb.InventoryReport) error`; a running ticker that reports metrics every 30s and static every 15 min. - -- [ ] **Step 1: Add client method** - -In `agent/internal/grpc/client.go`, mirroring `ReportUpdates`: - -```go -func (c *Client) ReportInventory(report *pb.InventoryReport) error { - ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second) - defer cancel() - _, err := c.client.ReportInventory(ctx, report) - return err -} -``` - -The report already carries `ServerId`/`AgentToken`; ensure the caller sets them (see Step 2). - -- [ ] **Step 2: Add the scheduler to the agent loop** - -Find where the agent starts its poll loop (the goroutine that calls `SyncKeys`/`ReportUpdates`). Add a parallel inventory ticker. `serverID`, `agentToken`, and the `*Client` are in scope there: - -```go - go func() { - tick := 0 - t := time.NewTicker(30 * time.Second) - defer t.Stop() - report := func(static bool) { - r := inventory.Collect(static) - r.ServerId = serverID - r.AgentToken = agentToken - if err := client.ReportInventory(r); err != nil { - log.Printf("report inventory: %v", err) - } - } - report(true) // send a full snapshot on startup - for range t.C { - tick++ - report(tick%30 == 0) // every 30th tick = 15 min → include static - } - }() -``` - -Add imports `"github.com/mrhid6/vantage/agent/internal/inventory"`, `time`, `log` if missing. Match variable names to the actual loop (e.g. the client may be named `c`). - -- [ ] **Step 3: Verify build** - -Run: `cd agent && go build ./... && go vet ./...` -Expected: success. - -- [ ] **Step 4: Commit** - -```bash -git add agent/internal/grpc/client.go agent/ -git commit -m "feat(agent): schedule inventory reporting (30s metrics, 15m static)" -``` - ---- - -## Task 5: Frontend — inventory panel on server detail - -**Files:** -- Modify: `web/lib/api.ts` (extend the `Server`/server-detail type with `inventory`) -- Modify: `web/app/servers/[id]/page.tsx` (add panel; enable polling) - -**Interfaces:** -- Consumes: server-detail query. - -- [ ] **Step 1: Add the inventory type** - -In `web/lib/api.ts`, add and attach to the server type used by the detail page: - -```ts -export interface Inventory { - cpu: { model?: string; cores?: number; usage_pct: number; load1?: number }; - memory: { total_bytes: number; used_bytes: number }; - swap_total_bytes: number; - swap_used_bytes: number; - partitions?: { device: string; mountpoint: string; fstype?: string; total_bytes: number; used_bytes: number }[]; - kernel?: string; - metrics_at?: string; - static_at?: string; -} -``` - -Add `inventory?: Inventory;` to the server detail interface. - -- [ ] **Step 2: Add a `formatBytes` helper + Inventory panel** - -In `web/app/servers/[id]/page.tsx`, add a helper and a panel component. Enable polling on the server-detail `useQuery` with `refetchInterval: 30000`. - -```tsx -function formatBytes(n: number): string { - if (!n) return "0 B"; - const u = ["B", "KB", "MB", "GB", "TB"]; - const i = Math.floor(Math.log(n) / Math.log(1024)); - return `${(n / Math.pow(1024, i)).toFixed(1)} ${u[i]}`; -} - -function UsageBar({ used, total }: { used: number; total: number }) { - const pct = total > 0 ? Math.min(100, (used / total) * 100) : 0; - return ( -
-
90 ? "bg-danger" : "bg-accent"}`} style={{ width: `${pct}%` }} /> -
- ); -} - -function InventoryPanel({ inv }: { inv: Inventory }) { - return ( - -

Inventory

-
-
-
CPU{inv.cpu.usage_pct.toFixed(0)}%
- -

{inv.cpu.model} · {inv.cpu.cores} cores · load {inv.cpu.load1?.toFixed(2)}

-
-
-
Memory{formatBytes(inv.memory.used_bytes)} / {formatBytes(inv.memory.total_bytes)}
- -
Swap{formatBytes(inv.swap_used_bytes)} / {formatBytes(inv.swap_total_bytes)}
- -
-
- {inv.partitions && inv.partitions.length > 0 && ( -
-

Partitions

-
- {inv.partitions.map((p) => ( -
-
- {p.mountpoint} - {formatBytes(p.used_bytes)} / {formatBytes(p.total_bytes)} · {p.fstype} -
- -
- ))} -
-
- )} - {inv.kernel &&

Kernel {inv.kernel}

} -
- ); -} -``` - -Render `{server.inventory && }` in the page body (ensure `Card`, `Inventory` are imported). Match how the page currently reads the server object. - -- [ ] **Step 3: Verify build** - -Run: `cd web && npm run build` -Expected: success. - -- [ ] **Step 4: Commit** - -```bash -git add web/lib/api.ts web/app/servers/[id]/page.tsx -git commit -m "feat(web): inventory panel on server detail" -``` - ---- - -## Task 6: End-to-end manual verification - -- [ ] **Step 1: Build all** - -Run: `cd server && go build ./... && cd ../agent && go build ./... && cd ../web && npm run build` -Expected: all succeed. - -- [ ] **Step 2: Smoke (if environment available)** - -With server + Mongo + a connected Linux agent: within ~30s the server detail page shows CPU %, RAM/swap bars; within 15 min (or on agent restart, which sends a full snapshot immediately) partitions, CPU model and kernel appear. Confirm metrics update roughly every 30s. - -- [ ] **Step 3: Commit any fixes** - -```bash -git add -A -git commit -m "fix: fleet inventory verification fixes" -``` - ---- - -# Service Monitoring (uptime-kuma replacement) - -Extends the fleet work: in-app service monitors replacing uptime-kuma. Monitors (HTTP/TCP/ICMP/TLS) run **server-side** (public endpoints) or **agent-side** (agent probes its own host). Both runners feed one server-side ingest pipeline: state → incidents → rollups → notifications. - -**Design:** validated in brainstorm 2026-07-21. Hybrid runners, all 4 check types, latest+incidents+rollups history, multi-channel notify (webhook/SMTP/Discord/Slack/Telegram), dedicated `SyncMonitors`/`ReportChecks` RPCs. - -**Build order — 3 phases, each shippable:** -- **P1 (Tasks 7–10):** data model, checker pkg, server scheduler, ingest pipeline, `/monitors` UI. Server-run only. No agent, no notify. -- **P2 (Tasks 11–12):** `SyncMonitors` + `ReportChecks` RPCs, agent checker + scheduler, agent-run monitors bound to a server. -- **P3 (Tasks 13–14):** notification channels + dispatch + settings UI. - -## Monitoring Global Constraints - -- Same as fleet: no tests this iteration; verify with `go build ./...`, `go vet ./...`, `npm run build`. JSON-codec gRPC — edit both pb files identically, mirror `ReportUpdates` wiring. Separate Go modules, so the checker pkg is **duplicated** in `server/` and `agent/` (same convention as pb files). -- Reuse existing patterns: REST handlers like `server/internal/api`, services like `server/internal/services/servers.go`, `db.Col(...)`, react-query + Tailwind UI like `web/app/servers`. - ---- - -## Task 7: Monitoring data model + checker package (server) - -**Files:** -- Create: `server/internal/models/monitor.go` -- Create: `server/internal/checker/checker.go` (+ `http.go`, `tcp.go`, `icmp.go`, `tls.go`) - -**Interfaces:** -- Produces: `models.Monitor` (+ `MonitorState`, `MonitorTarget`), `models.Incident`, `models.Rollup`. `checker.Run(ctx, models.Monitor) checker.Result` where `Result{Up bool; LatencyMs int; Message string; CertExpiry *time.Time}`. - -- [ ] **Step 1: Model** - -```go -type MonitorTarget struct { - URL string `bson:"url,omitempty" json:"url,omitempty"` - Host string `bson:"host,omitempty" json:"host,omitempty"` - Port int `bson:"port,omitempty" json:"port,omitempty"` - Method string `bson:"method,omitempty" json:"method,omitempty"` - ExpectedStatus int `bson:"expected_status,omitempty" json:"expected_status,omitempty"` - Keyword string `bson:"keyword,omitempty" json:"keyword,omitempty"` - TLSWarnDays int `bson:"tls_warn_days,omitempty" json:"tls_warn_days,omitempty"` -} -type MonitorState struct { - Status string `bson:"status" json:"status"` // up|down|pending - LastCheckAt *time.Time `bson:"last_check_at,omitempty" json:"last_check_at,omitempty"` - LatencyMs int `bson:"latency_ms" json:"latency_ms"` - Message string `bson:"message,omitempty" json:"message,omitempty"` - CertExpiryAt *time.Time `bson:"cert_expiry_at,omitempty" json:"cert_expiry_at,omitempty"` - Fails int `bson:"fails" json:"fails"` // consecutive failures -} -type Monitor struct { - MonitorID string `bson:"monitor_id" json:"monitor_id"` - Name string `bson:"name" json:"name"` - Type string `bson:"type" json:"type"` // http|tcp|icmp|tls - Target MonitorTarget `bson:"target" json:"target"` - IntervalSec int `bson:"interval_sec" json:"interval_sec"` - Runner string `bson:"runner" json:"runner"` // "server" or a server_id - Retries int `bson:"retries" json:"retries"` // consecutive fails before down - Enabled bool `bson:"enabled" json:"enabled"` - ChannelIDs []string `bson:"channel_ids,omitempty" json:"channel_ids,omitempty"` - State MonitorState `bson:"state" json:"state"` - CreatedAt time.Time `bson:"created_at" json:"created_at"` -} -type Incident struct { - IncidentID string `bson:"incident_id" json:"incident_id"` - MonitorID string `bson:"monitor_id" json:"monitor_id"` - StartedAt time.Time `bson:"started_at" json:"started_at"` - ResolvedAt *time.Time `bson:"resolved_at,omitempty" json:"resolved_at,omitempty"` - Cause string `bson:"cause,omitempty" json:"cause,omitempty"` -} -type Rollup struct { - MonitorID string `bson:"monitor_id" json:"monitor_id"` - PeriodStart time.Time `bson:"period_start" json:"period_start"` // hour bucket - Checks int `bson:"checks" json:"checks"` - UpCount int `bson:"up_count" json:"up_count"` - SumLatency int64 `bson:"sum_latency" json:"sum_latency"` -} -``` - -- [ ] **Step 2: Checker package** — `Run(ctx, m)` switches on `m.Type`: - - **http**: `http.Client` GET/HEAD `m.Target.URL`, assert status == ExpectedStatus (default 200), optional `Keyword` body contains; capture TLS peer cert expiry when https. - - **tcp**: `net.DialTimeout("tcp", host:port)`, latency = dial time. - - **icmp**: raw ICMP echo (agent/server run as root). Fall back to `net.Dial("ip4:icmp")`; on permission error return down with message. - - **tls**: `tls.Dial`, read `ConnectionState().PeerCertificates[0].NotAfter` → `CertExpiry`; down if within `TLSWarnDays` or expired. - - All: wrap with per-check timeout (min(IntervalSec, 10s)); `Result.Message` = short reason on failure. - -- [ ] **Step 3: Verify build** — `cd server && go build ./... && go vet ./...` - -- [ ] **Step 4: Commit** — `feat(server): monitor model + checker package` - ---- - -## Task 8: Ingest pipeline + rollups service - -**Files:** -- Create: `server/internal/services/monitors.go` - -**Interfaces:** -- Produces: `IngestResult(monitorID string, res checker.Result) error` — the single entry both runners use. `ListMonitors`, `GetMonitor`, `CreateMonitor`, `UpdateMonitor`, `DeleteMonitor`, `ListIncidents(monitorID)`, `UptimeRollups(monitorID, since)`. - -- [ ] **Step 1: `IngestResult`** — load monitor; compute new status with `Retries` threshold (increment `state.Fails` on failure, flip to `down` only when `Fails >= Retries`; reset + flip `up` on success). On **transition**: open incident (`down`) or resolve open incident (`up`), and enqueue notification (P3 — leave a `// TODO(P3): dispatch` hook now). Always `$set` state fields. Upsert current-hour `Rollup` (`$inc` checks/up_count/sum_latency). Use `db.Col("monitors")`, `db.Col("incidents")`, `db.Col("monitor_rollups")`, `context.WithTimeout`. - -- [ ] **Step 2: CRUD + queries** — standard service funcs mirroring `services/servers.go`. `UptimeRollups` aggregates buckets since a cutoff → uptime % + avg latency series. - -- [ ] **Step 3: Verify build** — `go build ./... && go vet ./...` - -- [ ] **Step 4: Commit** — `feat(server): monitor ingest pipeline, incidents, rollups` - ---- - -## Task 9: Server scheduler + REST API - -**Files:** -- Create: `server/internal/monitorsched/scheduler.go` -- Create: `server/internal/api/monitors.go` -- Modify: server bootstrap (wherever services/gRPC start) to launch the scheduler; router registration where `api` routes are mounted. - -**Interfaces:** -- Produces: a scheduler that ticks enabled `runner=="server"` monitors on their `IntervalSec` and calls `checker.Run` → `services.IngestResult`. REST: `GET/POST /api/monitors`, `GET/PUT/DELETE /api/monitors/:id`, `GET /api/monitors/:id/incidents`, `GET /api/monitors/:id/uptime`. - -- [ ] **Step 1: Scheduler** — on boot load monitors; per-monitor goroutine or a min-heap wheel keyed on next-run. Only `runner=="server"`. Reload on CRUD (simplest: re-read every N sec, or a reload channel fired by the service). Skip disabled. - -- [ ] **Step 2: REST handlers** — mirror an existing `server/internal/api` handler file for style + auth middleware. JSON in/out of `models.Monitor`. - -- [ ] **Step 3: Verify build** — `go build ./... && go vet ./...` - -- [ ] **Step 4: Commit** — `feat(server): server-run monitor scheduler + REST API` - ---- - -## Task 10: Frontend — monitors UI (P1) - -**Files:** -- Modify: `web/lib/api.ts` (Monitor types + bindings) -- Create: `web/app/monitors/page.tsx` (list), `web/app/monitors/[id]/page.tsx` (detail), `web/app/monitors/new/page.tsx` (create/edit form) -- Modify: main nav to add **Monitors** (same place Steps was added) - -**Interfaces:** -- Consumes: `/api/monitors*` (T9). - -- [ ] **Step 1: Types + api bindings** — `Monitor`, `MonitorState`, `Incident`, uptime series; `api.monitors.list/get/create/update/remove/incidents/uptime`. -- [ ] **Step 2: List page** — table: name, type, status badge (up/down/pending), uptime % (24h), latency, last check. `refetchInterval: 30000`. -- [ ] **Step 3: Detail page** — status header, heartbeat/uptime bars (24h + 30d from rollups), latency chart, incident timeline, cert expiry, assigned channels (read-only until P3). -- [ ] **Step 4: Create/edit form** — type-dependent fields (URL vs host/port), interval, retries, runner select (`server` or a registered server for agent-run — server option only wired in P2), enabled. -- [ ] **Step 5: Verify build** — `cd web && npm run build` -- [ ] **Step 6: Commit** — `feat(web): monitors list/detail/form UI` - ---- - -## Task 11: SyncMonitors + ReportChecks RPCs (P2) - -**Files:** -- Modify: `proto/vantage/v1/vantage.proto`, `server/internal/grpc/pb/vantage.pb.go`, `agent/internal/grpc/pb/vantage.pb.go`, `server/internal/grpc/server.go`, `agent/internal/grpc/client.go` - -**Interfaces:** -- Produces: `SyncMonitors(server_id, agent_token) -> repeated MonitorSpec`; `ReportChecks(server_id, agent_token, repeated CheckResult) -> ReportChecksResponse`. `MonitorSpec{monitor_id, type, target fields, interval_sec, retries}`. `CheckResult{monitor_id, up, latency_ms, message, cert_expiry_unix}`. - -- [ ] **Step 1: pb structs + proto** — add messages to both pb files + proto doc. -- [ ] **Step 2: Wire both RPCs** — mirror `ReportUpdates` plumbing (interface, Unimplemented stub, client method, `Vantage_ServiceDesc.Methods`, `_Vantage_*_Handler`) in both pb files. Server handlers on `vantageServer` (after `ReportUpdates` at server.go:78): `SyncMonitors` returns monitors where `runner==req.ServerId && enabled`; `ReportChecks` validates token then loops `services.IngestResult`. Client methods on `*Client` in client.go (after `ReportUpdates` at client.go:117). -- [ ] **Step 3: Verify build** — both modules `go build ./... && go vet ./...` -- [ ] **Step 4: Commit** — `feat(proto): SyncMonitors + ReportChecks RPCs` - ---- - -## Task 12: Agent checker + scheduler (P2) - -**Files:** -- Create: `agent/internal/checker/` (duplicate of server checker pkg) -- Create: `agent/internal/monitors/monitors.go` (poll + run + report loop) -- Modify: agent main loop to start it (alongside the sync loop in `agent/internal/sync` / the inventory ticker from Task 4) - -**Interfaces:** -- Consumes: `client.SyncMonitors`, `client.ReportChecks`, agent `checker`. - -- [ ] **Step 1: Duplicate checker pkg** into agent module (identical logic; imports agent pb). -- [ ] **Step 2: Monitor loop** — poll `SyncMonitors` every 30s for assigned specs; per-spec ticker on `IntervalSec` runs `checker.Run`; batch `CheckResult`s and `ReportChecks`. `serverID`/`agentToken`/`*Client` in scope from the existing loop. -- [ ] **Step 3: Verify build** — `cd agent && go build ./... && go vet ./...` (+ `GOOS=windows go build ./...`; icmp may no-op on Windows). -- [ ] **Step 4: Commit** — `feat(agent): agent-run monitor scheduler` - ---- - -## Task 13: Notification channels + dispatch (P3) - -**Files:** -- Create: `server/internal/models/channel.go`, `server/internal/services/channels.go`, `server/internal/notify/` (`dispatch.go`, `webhook.go`, `smtp.go`, `discord.go`, `slack.go`, `telegram.go`), `server/internal/api/channels.go` -- Modify: `server/internal/services/monitors.go` (replace the P2 `// TODO(P3): dispatch` hook) - -**Interfaces:** -- Produces: `models.NotificationChannel{channel_id, name, type, config map, enabled}`. `notify.Dispatch(channel, event)` where `event` = monitor + old/new status + message. `notify.Test(channel)`. - -- [ ] **Step 1: Model + CRUD service + REST** (`/api/channels*`, incl. `POST /api/channels/:id/test`). -- [ ] **Step 2: Dispatch abstraction** — webhook/discord/slack/telegram are HTTP POST with per-type JSON payload; SMTP via `net/smtp`. Per-monitor routing via `monitor.ChannelIDs`; resend interval so an ongoing `down` re-alerts at most every N min (track `last_notified_at` on monitor state). -- [ ] **Step 3: Fire on transition** — in `IngestResult`, on up/down flip resolve channels and `notify.Dispatch` each (goroutine, best-effort, log failures). -- [ ] **Step 4: Verify build** — `go build ./... && go vet ./...` -- [ ] **Step 5: Commit** — `feat(server): multi-channel monitor notifications` - ---- - -## Task 14: Frontend — notification settings (P3) - -**Files:** -- Modify: `web/lib/api.ts` (channel types + bindings), `web/app/settings/` (add notifications section/page) -- Modify: monitor create/edit form (Task 10) to select channels - -**Interfaces:** -- Consumes: `/api/channels*`. - -- [ ] **Step 1: Channel types + api bindings.** -- [ ] **Step 2: Settings UI** — list/add/edit channels, type-dependent config fields, **Test** button hitting `/api/channels/:id/test`. -- [ ] **Step 3: Wire channel multi-select** into the monitor form. -- [ ] **Step 4: Verify build** — `cd web && npm run build` -- [ ] **Step 5: Commit** — `feat(web): notification channel settings UI` - ---- - -## Self-Review Notes - -- **Spec coverage:** §3 model → T1; §4 RPC → T1; §5 collectors + scheduler → T3, T4; §6 handler/store → T2; §7 frontend → T5. Split cadence (30s metrics / 15m static) in T4 scheduler; merge rules preserving static in T2 `StoreInventory`. Tests omitted per Global Constraints. -- **Startup snapshot:** agent sends `Collect(true)` immediately so static fields populate without waiting 15 min. -- **Types consistent:** `InventoryReport` field names identical across proto, both pb files, store service, and TS interface (`usage_pct`, `used_bytes`, `total_bytes`, `swap_*`). -- **Follow-ups (out of scope):** time-series history, usage alerting, Windows collectors, servers-list CPU/RAM badges. -- **Monitoring (Tasks 7–14):** hybrid runner service-monitor replacing uptime-kuma, added 2026-07-21. 3 phases — P1 server-run engine+UI (T7–10), P2 agent-run RPCs (T11–12), P3 multi-channel notify (T13–14). Single `IngestResult` pipeline for both runners; checker pkg duplicated per module (pb convention). Design: brainstorm 2026-07-21. Follow-ups out of scope: status pages, maintenance windows, per-check auth headers, ICMP on Windows. diff --git a/docs/superpowers/specs/2026-07-20-fleet-inventory-design.md b/docs/superpowers/specs/2026-07-20-fleet-inventory-design.md deleted file mode 100644 index b449c7a..0000000 --- a/docs/superpowers/specs/2026-07-20-fleet-inventory-design.md +++ /dev/null @@ -1,142 +0,0 @@ -# Fleet Inventory — Design - -**Date:** 2026-07-20 -**Status:** Approved (design) — ready for implementation planning -**Scope:** Fleet Inventory only. Server Workflows and SaaS/auth are separate sub-projects. - ---- - -## 1. Summary - -Each agent collects hardware/OS inventory about its host and reports it to the server, which stores the latest snapshot per server and surfaces it in the UI. Two cadences: - -- **Metrics (near-real-time):** CPU load/usage, RAM used/total, swap used/total — every **30s** (aligned with existing poll rhythm). -- **Static inventory (slow):** disks, partitions and their usage, CPU model/cores, total RAM, OS details — every **15 min**. - -Transport: a **new unary gRPC `ReportInventory` RPC** (mirrors the existing `ReportUpdates` pattern). No streaming. - ---- - -## 2. Locked decisions - -| Topic | Decision | -|-------|----------| -| Transport | New `ReportInventory` unary RPC. | -| Cadence | Metrics every 30s; static inventory every 15 min. One RPC carries both, but static fields are only populated on the 15-min tick (empty/omitted otherwise → server keeps prior static snapshot). | -| Storage | Latest snapshot embedded on the `servers` document (`inventory` sub-doc). No history/time-series in v1. | -| Collection | Pure-Go where practical (`/proc`, `gopsutil`-style). Agent already runs as root. | -| Platform | Linux primary; Windows agent populates what it can, leaves the rest empty. | - ---- - -## 3. Data model - -Add an `Inventory` sub-document to the existing `Server` model (`server/internal/models/server.go`): - -```go -type CPUInfo struct { - Model string `bson:"model,omitempty" json:"model,omitempty"` - Cores int `bson:"cores,omitempty" json:"cores,omitempty"` - UsagePct float64 `bson:"usage_pct" json:"usage_pct"` // metrics tick - Load1 float64 `bson:"load1,omitempty" json:"load1,omitempty"` -} - -type MemInfo struct { - TotalBytes uint64 `bson:"total_bytes" json:"total_bytes"` - UsedBytes uint64 `bson:"used_bytes" json:"used_bytes"` // metrics tick -} - -type Partition struct { - Device string `bson:"device" json:"device"` - Mountpoint string `bson:"mountpoint" json:"mountpoint"` - Fstype string `bson:"fstype,omitempty" json:"fstype,omitempty"` - TotalBytes uint64 `bson:"total_bytes" json:"total_bytes"` - UsedBytes uint64 `bson:"used_bytes" json:"used_bytes"` -} - -type Inventory struct { - CPU CPUInfo `bson:"cpu" json:"cpu"` - Memory MemInfo `bson:"memory" json:"memory"` - SwapTotalBytes uint64 `bson:"swap_total_bytes" json:"swap_total_bytes"` - SwapUsedBytes uint64 `bson:"swap_used_bytes" json:"swap_used_bytes"` - Partitions []Partition `bson:"partitions,omitempty" json:"partitions,omitempty"` - Kernel string `bson:"kernel,omitempty" json:"kernel,omitempty"` - MetricsAt *time.Time `bson:"metrics_at,omitempty" json:"metrics_at,omitempty"` - StaticAt *time.Time `bson:"static_at,omitempty" json:"static_at,omitempty"` -} -``` - -Add `Inventory *Inventory` field to `Server`. - -Server-side update rules: -- Metrics fields (`cpu.usage_pct`, `cpu.load1`, `memory.used_bytes`, swap used) always updated + `metrics_at`. -- Static fields (`cpu.model/cores`, `memory.total_bytes`, `partitions`, `kernel`, swap total) updated only when the report includes them (non-zero/non-empty) + `static_at`. - ---- - -## 4. gRPC protocol (`proto/vantage/v1/vantage.proto` + both `pb.go` files) - -```protobuf -rpc ReportInventory(InventoryReport) returns (InventoryReportResponse); - -message InventoryReport { - string server_id = 1; - string agent_token = 2; - bool include_static = 3; // true on the 15-min tick - CPUReport cpu = 4; - MemReport memory = 5; - uint64 swap_total = 6; - uint64 swap_used = 7; - repeated PartitionReport partitions = 8; // only when include_static - string kernel = 9; // only when include_static -} -message CPUReport { string model = 1; int32 cores = 2; double usage_pct = 3; double load1 = 4; } -message MemReport { uint64 total_bytes = 1; uint64 used_bytes = 2; } -message PartitionReport { string device = 1; string mountpoint = 2; string fstype = 3; uint64 total_bytes = 4; uint64 used_bytes = 5; } -message InventoryReportResponse {} -``` - -Hand-written JSON-codec structs added to `server/internal/grpc/pb/vantage.pb.go` and `agent/internal/grpc/pb/vantage.pb.go`, plus the RPC method wiring (service interface, client method, handler registration) mirroring `ReportUpdates`. - ---- - -## 5. Agent collection (`agent/internal/inventory/`) - -- `Collect(includeStatic bool) *pb.InventoryReport` — reads: - - CPU usage: sample `/proc/stat` delta; load from `/proc/loadavg`; model/cores from `/proc/cpuinfo` (static). - - Memory/swap: `/proc/meminfo`. - - Partitions: `/proc/mounts` filtered to real filesystems + `statfs` for total/used (static). - - Kernel: `uname` / `/proc/version` (static). - - Windows: best-effort via `wmic`/PS or leave empty. -- Scheduler in the agent main loop: a 30s ticker calls `Collect(false)` and `ReportInventory`; every 30th tick (15 min) calls `Collect(true)`. -- Reuse existing gRPC client; add `Client.ReportInventory(...)` like `ReportUpdates`. - -Prefer implementing the `/proc` readers directly (no new heavy deps) unless a `gopsutil` dependency is already vendored. - ---- - -## 6. Server handler + service - -- gRPC handler `ReportInventory` in `server/internal/grpc/server.go`: validate agent token (`ValidateAgentToken`), then call `services.StoreInventory(serverID, report)`. -- `services.StoreInventory` (in `server/internal/services/inventory.go`): builds the `$set` per the update rules in §3 and `UpdateOne` on `servers`. - ---- - -## 7. Frontend - -Surface inventory on the existing server detail page (`web/app/servers/[id]/page.tsx`) — add an "Inventory" panel: -- CPU usage gauge + model/cores, load. -- RAM used/total bar, swap bar. -- Partitions table: device, mount, fstype, used/total with a usage bar. -- "Updated Xs ago" from `metrics_at`/`static_at`. - -Optionally add compact CPU/RAM badges to the servers list (`web/app/servers/page.tsx`). Reuse `@/components/ui` + Tailwind tokens. Poll the server detail query while the page is open (react-query `refetchInterval` ~30s) so metrics stay fresh. - ---- - -## 8. Out of scope - -- Time-series history / graphs (only latest snapshot stored). -- Alerting thresholds on usage (settings/alerts is a separate concern). -- Per-process / network / GPU inventory. -- Tests (skipped, consistent with the Workflows iteration).