Compare commits

..
36 Commits
Author SHA1 Message Date
mrhid6 71240f183c fix: Fixes to server shutdown stream
Chart Release / chart (push) Successful in 21s
Server Deploy / deploy (push) Successful in 1m2s
Agent Release / build (push) Successful in 43s
Agent Release / msi (push) Successful in 49s
2026-07-31 16:44:47 +01:00
mrhid6 01e8b0ba44 feat: Better debugging for console
Chart Release / chart (push) Successful in 11s
Server Deploy / deploy (push) Successful in 1m25s
2026-07-31 16:31:19 +01:00
mrhid6 2aa4784518 feat: Better debugging for console
Chart Release / chart (push) Successful in 18s
Server Deploy / deploy (push) Successful in 59s
2026-07-31 16:13:51 +01:00
mrhid6 f611cae438 feat: Better debugging for console
Chart Release / chart (push) Successful in 11s
Server Deploy / deploy (push) Successful in 1m22s
2026-07-31 16:00:58 +01:00
mrhid6 1eb98ef962 feat: Better debugging for console
Chart Release / chart (push) Successful in 12s
Server Deploy / deploy (push) Successful in 1m22s
2026-07-31 15:51:26 +01:00
mrhid6 6f86496f10 fix: Ffixes to console
Chart Release / chart (push) Successful in 18s
Server Deploy / deploy (push) Successful in 1m9s
2026-07-31 15:05:21 +01:00
mrhid6 57a9b18102 fix: Guacd connection ip
Server Deploy / deploy (push) Successful in 9s
Chart Release / chart (push) Successful in 11s
2026-07-31 14:52:20 +01:00
mrhid6 36995fa62b fix: Fixed install and update scripts
Chart Release / chart (push) Successful in 9s
Server Deploy / deploy (push) Successful in 1m20s
2026-07-31 12:10:32 +01:00
mrhid6 9121fc461f fix: Fixed chart api routes for update
Server Deploy / deploy (push) Successful in 15s
Chart Release / chart (push) Successful in 10s
2026-07-31 12:03:45 +01:00
mrhid6 fc56bae5f9 chore: Bump chart version
Chart Release / chart (push) Successful in 18s
Server Deploy / deploy (push) Successful in 8s
Agent Release / build (push) Successful in 38s
Agent Release / msi (push) Successful in 57s
2026-07-31 11:53:08 +01:00
mrhid6 ac75b3ef76 feat: chart deployment Type added
Chart Release / chart (push) Successful in 20s
Server Deploy / deploy (push) Successful in 35s
2026-07-31 11:52:31 +01:00
mrhid6 e6fe463216 feat: Updated for api ingress routes
Chart Release / chart (push) Successful in 25s
Server Deploy / deploy (push) Successful in 4m26s
2026-07-31 11:21:15 +01:00
mrhid6 8528f14ed7 feat: Added ingress to chart
Chart Release / chart (push) Successful in 10s
Server Deploy / deploy (push) Successful in 1m24s
2026-07-31 10:49:20 +01:00
mrhid6 df1d9658f5 fix: Chart build
Chart Release / chart (push) Successful in 10s
Server Deploy / deploy (push) Successful in 1m20s
2026-07-31 10:41:41 +01:00
mrhid6 9f9b384481 fix: Fixed chart version
Chart Release / chart (push) Failing after 13s
2026-07-31 10:37:04 +01:00
mrhid6 165114471f fix: Fixes to running on kubernetes
Chart Release / chart (push) Failing after 13s
Server Deploy / deploy (push) Successful in 6m35s
2026-07-31 10:34:10 +01:00
mrhid6 de78688093 feat: authenticate the server's Redis connection
InitRedis now takes a username and password, read from REDIS_USERNAME and
REDIS_PASSWORD, matching what admin has always done. Both empty keeps an
unauthenticated Redis working; a password with an empty username is what a
legacy requirepass instance needs, since go-redis then sends AUTH with one
argument instead of two.

This is what lets a Kubernetes install point at a managed Redis instead of
the bundled one.
2026-07-31 09:36:31 +01:00
mrhid6 bbf9f72fd3 feat: Docker and helm charts
Server Deploy / deploy (push) Successful in 5m26s
Agent Release / build (push) Successful in 10m45s
Agent Release / msi (push) Successful in 1m31s
2026-07-31 09:28:54 +01:00
mrhid6 978b665aa6 fix: stop local relay teardown from logging a spurious proxy_failed reason
Session.Close now closing its own accepted conn (from the prior fix wave)
made net.ErrClosed on the guacd-side reader indistinguishable from a real
remote failure, so a normal browser-tab close could race the handler's
defer and intermittently log console.proxy_failed on a healthy session.
Add a closing flag, set before Close's sync.Once body actually tears
anything down, that setReason respects -- a deliberate local teardown can
no longer produce or race in a failure reason, while Close's own explicit
reason argument still wins normally.
2026-07-31 09:25:27 +01:00
mrhid6 1fe608f531 fix: bound and complete console relay teardown, restore proxy_failed audit
- Arm the unclaimed-relay watchdog in NewSession rather than Serve, so an
  agent that never opens its ProxyStream is bounded to 10s and reports
  reason "agent_timeout", per the design spec's failure-mode table.
- Session.Close now also closes the accepted net.Conn (stored via setConn),
  so ConsoleProxy.Close() is an unconditional kill of the whole relay chain
  instead of only closing an already-idle listener.
- Emit console.proxy_failed and end the console session from a defer in
  consoleTunnel guarded on relay.Reason(), since guac's OnDisconnect never
  runs when the connect callback errors -- which is the path every relay
  failure this feature introduces takes. Update the two docsite
  troubleshooting rows to match what the audit event can now actually show.
2026-07-31 09:21:07 +01:00
mrhid6 1e1546cb60 docs: document the agent-relayed console proxy
Every console session now rides the agent's outbound gRPC connection
instead of a direct guacd-to-target dial, so it works for servers
behind NAT and now requires a live agent (409 agent_offline
otherwise). Documents PROXY_ADVERTISE_HOST / PROXY_LISTEN_HOST and
corrects reachability claims across the docsite and CLAUDE.md.
2026-07-31 09:10:05 +01:00
mrhid6 119d8694d1 feat: Reap admin free instance license 2026-07-30 14:42:06 +01:00
mrhid6 8d43c689f5 feat: Reap admin free instance 2026-07-30 14:31:43 +01:00
mrhid6 05f10ed3c9 feat: record relay proxy_id and port in console audit events 2026-07-29 13:10:58 +01:00
mrhid6 c0bec3737b feat: route every console session through the agent relay 2026-07-29 13:07:30 +01:00
mrhid6 59d147fe4d feat: handle OpenProxyCmd in the agent command stream 2026-07-29 13:03:02 +01:00
mrhid6 9e38a01e3d feat: add agent-side console relay 2026-07-29 12:59:24 +01:00
mrhid6 20a302f84a feat: add OpenConsoleProxy service facade 2026-07-29 12:55:36 +01:00
mrhid6 ba2e263d00 fix: collapse ProxyStream auth failures into one indistinguishable response 2026-07-29 12:52:53 +01:00
mrhid6 a000703199 feat: add ProxyStream handler with scoped single-use auth 2026-07-29 12:50:20 +01:00
mrhid6 8fcda63742 fix: avoid closing proxy relay listener before validating remote source 2026-07-29 12:47:35 +01:00
mrhid6 3363ac9dad feat: add console proxy session relay 2026-07-29 12:43:04 +01:00
mrhid6 a7e338b171 feat: add console proxy session registry 2026-07-29 12:40:22 +01:00
mrhid6 bc79daab48 feat: add ProxyStream wire types for agent-relayed console 2026-07-29 12:37:37 +01:00
mrhid6 d3d8dba3ff docs: Implementation plan for agent-relayed console proxy 2026-07-29 12:26:16 +01:00
mrhid6 6d047e25ab docs: Design for agent-relayed console proxy 2026-07-29 12:16:45 +01:00
57 changed files with 6674 additions and 336 deletions
+221
View File
@@ -0,0 +1,221 @@
name: Chart Release
on:
# Every push that touches the chart is validated. Publishing is separate and
# deliberate: a chart version is immutable in the registry once pushed, so
# it must come from a tag someone chose, not from whatever landed on main.
# No `paths` filter on push, deliberately. A paths filter applies to tag
# pushes too, so tagging a commit that happened not to touch the chart
# would skip the publish entirely — a release that silently does nothing.
# Validation is seconds of helm rendering; running it on every push to main
# is cheaper than that failure mode.
push:
branches:
- main
tags:
- "chart/v*"
pull_request:
paths:
- "deploy/chart/**"
workflow_dispatch:
env:
CHART_DIR: deploy/chart/vantage
HELM_VERSION: v3.16.3
jobs:
chart:
runs-on: ubuntu-docker
container: alpine:3.21
steps:
# git for actions/checkout, curl for both the Helm download and the
# registry upload, tar because the Helm tarball is not self-extracting.
- name: Setup
run: apk add --no-cache bash curl git tar nodejs npm
- name: Install Helm
run: |
set -eu
curl -fsSL "https://get.helm.sh/helm-${HELM_VERSION}-linux-amd64.tar.gz" \
| tar -xz -C /tmp linux-amd64/helm
mv /tmp/linux-amd64/helm /usr/local/bin/helm
helm version --short
- name: Checkout
uses: actions/checkout@v4
- name: Lint
run: helm lint "$CHART_DIR"
# Rendering is the real test. `helm lint` accepts a chart whose
# templates fail to execute, and every guard in this chart is a
# template `fail` that only fires during rendering.
- name: Render default values
run: helm template test "$CHART_DIR" > /dev/null
- name: Render a multi-replica install
run: |
helm template test "$CHART_DIR" \
--set server.replicaCount=3 \
--set web.replicaCount=3 > /dev/null
# The reaper deletes whole instances, so "does this env appear only
# in cloud mode" is worth asserting rather than eyeballing.
- name: Check the reaper is cloud-only
run: |
set -eu
if helm template test "$CHART_DIR" | grep -q FREE_INSTANCE_REAP_AFTER; then
echo "FREE_INSTANCE_REAP_AFTER is set on a self-hosted render"
exit 1
fi
if ! helm template test "$CHART_DIR" \
--set server.env.deploymentType=cloud \
| grep -q FREE_INSTANCE_REAP_AFTER; then
echo "FREE_INSTANCE_REAP_AFTER is missing from a cloud render"
exit 1
fi
echo "ok: reaper configured in cloud mode only"
- name: Render against external Redis and MongoDB
run: |
helm template test "$CHART_DIR" \
--set redis.enabled=false \
--set redis.addr=redis.example.com:6379 \
--set mongo.enabled=false \
--set server.env.mongoUri=mongodb://mongo.example.com:27017/vantage > /dev/null
- name: Render with the Traefik ingress
run: |
helm template test "$CHART_DIR" \
--set ingress.enabled=true \
--set ingress.web.host=vantage.example.com \
--set ingress.grpc.host=agents.example.com \
--set ingress.tls.certResolver=letsencrypt \
--set server.env.grpcHost=agents.example.com:443 > /dev/null
# The shape the cloud deployment actually uses: a wildcard tenant
# namespace, /api and /auth routed at the edge, and no apex — that
# belongs to the marketing site, which this chart does not deploy.
- name: Render a wildcard host with edge-routed API paths
run: |
helm template test "$CHART_DIR" \
--set ingress.enabled=true \
--set 'ingress.web.host=*.vantage.example.com' \
--set ingress.api.enabled=true \
--set ingress.grpc.host=agents.example.com \
--set server.env.grpcHost=agents.example.com:443 \
--set ingress.tls.secretName=vantage-tls \
--set ingress.tls.grpcSecretName=agents-tls > /dev/null
# The guards are load-bearing, so their absence is a regression the
# same way a broken render is. Each of these must fail.
- name: Check the guards still refuse bad values
run: |
set -eu
refuses() {
desc="$1"; shift
if helm template test "$CHART_DIR" "$@" > /dev/null 2>&1; then
echo "GUARD MISSING: $desc was accepted"
exit 1
fi
echo "ok: refused $desc"
}
refuses "mongo disabled with an in-chart URI" \
--set mongo.enabled=false
refuses "redis disabled with no external address" \
--set redis.enabled=false
refuses "multiple replicas on a ReadWriteOnce volume" \
--set server.replicaCount=2 --set server.persistence.enabled=true
refuses "ingress with no web host" \
--set ingress.enabled=true
refuses "edge-routed API with an empty path list" \
--set ingress.enabled=true \
--set ingress.web.host=vantage.example.com \
--set ingress.grpc.enabled=false \
--set ingress.api.enabled=true \
--set 'ingress.api.paths=null'
refuses "gRPC ingress with no host" \
--set ingress.enabled=true \
--set ingress.web.host=vantage.example.com \
--set server.env.grpcHost=agents.example.com:443
refuses "gRPC ingress while grpcHost is still in-cluster" \
--set ingress.enabled=true \
--set ingress.web.host=vantage.example.com \
--set ingress.grpc.host=agents.example.com
- name: Read the chart version
id: chart
run: |
set -eu
VERSION="$(grep '^version:' "$CHART_DIR/Chart.yaml" | awk '{print $2}')"
echo "version=$VERSION" >> "$GITHUB_OUTPUT"
echo "chart version is $VERSION"
# Chart.yaml is the source of truth for the version; the tag only
# says "publish this one". A mismatch is a mistake worth stopping
# for — the alternative is stamping the tag over Chart.yaml, which
# leaves the repository disagreeing with what was published.
- name: Check the tag matches Chart.yaml
if: startsWith(github.ref, 'refs/tags/chart/v')
run: |
set -eu
TAG_VERSION="${GITHUB_REF_NAME#chart/v}"
CHART_VERSION="${{ steps.chart.outputs.version }}"
if [ "$TAG_VERSION" != "$CHART_VERSION" ]; then
echo "tag chart/v$TAG_VERSION does not match Chart.yaml version $CHART_VERSION"
echo "bump version: in $CHART_DIR/Chart.yaml, or retag."
exit 1
fi
- name: Package
run: |
set -eu
mkdir -p dist
helm package "$CHART_DIR" --destination dist
ls -l dist
- name: Publish to the Gitea chart registry
if: startsWith(github.ref, 'refs/tags/chart/v')
env:
# github.server_url is this Gitea instance, so the registry
# host needs no variable of its own and cannot drift from it.
REGISTRY: ${{ github.server_url }}/api/packages/${{ github.repository_owner }}/helm/api/charts
# The same pair server-deploy.yml uses for `docker login`.
# RELEASE_TOKEN, not REGISTRY_PASSWORD: the latter is named in
# the docs but set by no workflow, and an unset secret becomes
# an empty password, which Gitea reports as "Failed to
# authenticate user" rather than as a missing credential.
REGISTRY_USER: ${{ secrets.REGISTRY_USER }}
REGISTRY_TOKEN: ${{ secrets.RELEASE_TOKEN }}
CHART_VERSION: ${{ steps.chart.outputs.version }}
run: |
set -eu
PKG="dist/vantage-${CHART_VERSION}.tgz"
test -f "$PKG"
# Checked explicitly, because the failure it prevents is a
# 401 that looks like a permissions problem on the token that
# was never sent.
if [ -z "${REGISTRY_USER}" ] || [ -z "${REGISTRY_TOKEN}" ]; then
echo "REGISTRY_USER or RELEASE_TOKEN is not set on this repository."
echo "RELEASE_TOKEN needs the write:package scope to publish a chart."
exit 1
fi
echo "publishing to ${REGISTRY} as ${REGISTRY_USER}"
# --fail-with-body so an HTTP error is a failed step with the
# server's explanation, rather than a green run that published
# nothing. A repeated version is rejected by the registry;
# that is the intended behaviour, not something to retry past.
curl --fail-with-body -sS \
--user "${REGISTRY_USER}:${REGISTRY_TOKEN}" \
-X POST \
--upload-file "$PKG" \
"$REGISTRY"
echo "published vantage ${CHART_VERSION}"
echo " helm repo add vantage ${{ github.server_url }}/api/packages/${{ github.repository_owner }}/helm"
echo " helm install vantage vantage/vantage --version ${CHART_VERSION}"
+3 -1
View File
@@ -3,6 +3,7 @@ dist
build
.env
.env.bck
.env.live
docs/*
!docs/superpowers/
.superpowers
@@ -12,4 +13,5 @@ installer/nssm.zip
installer/checksums-msi.txt
.next
*.tsbuildinfo
graphify-out
graphify-out
docker-compose.live.yml
+10
View File
@@ -77,6 +77,16 @@ func Run(ctx context.Context) error {
}
}
if lic.ExpiresAt.Add(reapAfter).Before(now) {
if _, err := db.Admin("admin_instances").DeleteOne(ctx, bson.M{"instance_id": inst.InstanceID}); err != nil {
log.Printf("lifecycle: delete instance %s: %v", inst.InstanceID, err)
}
if _, err := db.Admin("licenses").DeleteMany(ctx, bson.M{"instance_id": inst.InstanceID}); err != nil {
log.Printf("lifecycle: delete licenses for instance %s: %v", inst.InstanceID, err)
}
}
due := dueNotice(now, lic.ExpiresAt, inst.NoticesSent)
if due == "" {
continue
+4
View File
@@ -151,3 +151,7 @@ func (c *Client) ReportChecks(serverID, agentToken string, results []pb.CheckRes
func (c *Client) CommandStream(ctx context.Context) (pb.Vantage_CommandStreamClient, error) {
return c.client.CommandStream(ctx)
}
func (c *Client) ProxyStream(ctx context.Context) (pb.Vantage_ProxyStreamClient, error) {
return c.client.ProxyStream(ctx)
}
+82
View File
@@ -131,6 +131,32 @@ type ReportChecksResponse struct{}
type ApplyUpdatesCmd struct{}
type OpenProxyCmd struct {
ProxyId string `json:"proxy_id"`
Port uint32 `json:"port"`
}
type ProxyOpen struct {
ServerId string `json:"server_id"`
AgentToken string `json:"agent_token"`
ProxyId string `json:"proxy_id"`
}
type ProxyClose struct {
Reason string `json:"reason,omitempty"`
}
type ProxyClientMsg struct {
Open *ProxyOpen `json:"open,omitempty"`
Data []byte `json:"data,omitempty"`
Close *ProxyClose `json:"close,omitempty"`
}
type ProxyServerMsg struct {
Data []byte `json:"data,omitempty"`
Close *ProxyClose `json:"close,omitempty"`
}
type ServerCommand struct {
CommandId string `json:"command_id"`
GenerateKey *GenerateKeyCmd `json:"generate_key,omitempty"`
@@ -139,6 +165,7 @@ type ServerCommand struct {
ApplyUpdates *ApplyUpdatesCmd `json:"apply_updates,omitempty"`
RunStep *RunStepCmd `json:"run_step,omitempty"`
CleanupWorkspace *CleanupWorkspaceCmd `json:"cleanup_workspace,omitempty"`
OpenProxy *OpenProxyCmd `json:"open_proxy,omitempty"`
}
@@ -254,6 +281,51 @@ func (s *keyManagerCommandStreamServer) Recv() (*AgentMessage, error) {
return m, nil
}
type Vantage_ProxyStreamServer interface {
Send(*ProxyServerMsg) error
Recv() (*ProxyClientMsg, error)
grpc.ServerStream
}
type vantageProxyStreamServer struct {
grpc.ServerStream
}
func (s *vantageProxyStreamServer) Send(m *ProxyServerMsg) error {
return s.ServerStream.SendMsg(m)
}
func (s *vantageProxyStreamServer) Recv() (*ProxyClientMsg, error) {
m := new(ProxyClientMsg)
if err := s.ServerStream.RecvMsg(m); err != nil {
return nil, err
}
return m, nil
}
type Vantage_ProxyStreamClient interface {
Send(*ProxyClientMsg) error
Recv() (*ProxyServerMsg, error)
CloseSend() error
grpc.ClientStream
}
type vantageProxyStreamClient struct {
grpc.ClientStream
}
func (c *vantageProxyStreamClient) Send(m *ProxyClientMsg) error {
return c.ClientStream.SendMsg(m)
}
func (c *vantageProxyStreamClient) Recv() (*ProxyServerMsg, error) {
m := new(ProxyServerMsg)
if err := c.ClientStream.RecvMsg(m); err != nil {
return nil, err
}
return m, nil
}
type VantageClient interface {
Register(ctx context.Context, in *RegisterRequest, opts ...grpc.CallOption) (*RegisterResponse, error)
SyncKeys(ctx context.Context, in *SyncRequest, opts ...grpc.CallOption) (*SyncResponse, error)
@@ -263,6 +335,7 @@ type VantageClient interface {
SyncMonitors(ctx context.Context, in *SyncMonitorsRequest, opts ...grpc.CallOption) (*SyncMonitorsResponse, error)
ReportChecks(ctx context.Context, in *ReportChecksRequest, opts ...grpc.CallOption) (*ReportChecksResponse, error)
CommandStream(ctx context.Context, opts ...grpc.CallOption) (Vantage_CommandStreamClient, error)
ProxyStream(ctx context.Context, opts ...grpc.CallOption) (Vantage_ProxyStreamClient, error)
}
type UnimplementedVantageServer struct{}
@@ -349,3 +422,12 @@ func (c *keyManagerClient) CommandStream(ctx context.Context, opts ...grpc.CallO
}
return &vantageCommandStreamClient{stream}, nil
}
func (c *keyManagerClient) ProxyStream(ctx context.Context, opts ...grpc.CallOption) (Vantage_ProxyStreamClient, error) {
desc := &grpc.StreamDesc{StreamName: "ProxyStream", ServerStreams: true, ClientStreams: true}
stream, err := c.cc.NewStream(ctx, desc, "/vantage.v1.Vantage/ProxyStream", opts...)
if err != nil {
return nil, err
}
return &vantageProxyStreamClient{stream}, nil
}
+114
View File
@@ -0,0 +1,114 @@
// Package agentproxy relays a single TCP connection between a local service and
// the control plane, so a control plane that cannot route to this host's network
// can still open a console session.
//
// The dial host is hardcoded to loopback. The control plane supplies only a
// port, and nothing in this package can be made to dial anywhere else.
package agentproxy
import (
"errors"
"fmt"
"io"
"net"
"strconv"
"time"
"gitea.hostxtra.co.uk/mrhid6/vantage/agent/internal/grpc/pb"
)
const (
loopbackHost = "127.0.0.1"
chunkSize = 32 * 1024
dialTimeout = 10 * time.Second
)
// Stream is the agent's half of a ProxyStream.
type Stream interface {
Send(*pb.ProxyClientMsg) error
Recv() (*pb.ProxyServerMsg, error)
CloseSend() error
}
// Open dials the local port, announces itself on the stream, and relays until
// either side ends. A refused dial is reported as an explicit close so the
// operator sees a reason rather than a hang.
func Open(stream Stream, serverID, agentToken, proxyID string, port uint32) error {
conn, dialErr := net.DialTimeout("tcp",
net.JoinHostPort(loopbackHost, strconv.Itoa(int(port))), dialTimeout)
if err := stream.Send(&pb.ProxyClientMsg{Open: &pb.ProxyOpen{
ServerId: serverID,
AgentToken: agentToken,
ProxyId: proxyID,
}}); err != nil {
if conn != nil {
_ = conn.Close()
}
return fmt.Errorf("send open: %w", err)
}
if dialErr != nil {
_ = stream.Send(&pb.ProxyClientMsg{Close: &pb.ProxyClose{
Reason: "dial_refused: " + dialErr.Error(),
}})
_ = stream.CloseSend()
return fmt.Errorf("dial 127.0.0.1:%d: %w", port, dialErr)
}
defer conn.Close()
return relay(conn, stream)
}
func relay(conn net.Conn, stream Stream) error {
errCh := make(chan error, 2)
// local service -> control plane
go func() {
buf := make([]byte, chunkSize)
for {
n, err := conn.Read(buf)
if n > 0 {
chunk := make([]byte, n)
copy(chunk, buf[:n])
if sendErr := stream.Send(&pb.ProxyClientMsg{Data: chunk}); sendErr != nil {
errCh <- sendErr
return
}
}
if err != nil {
errCh <- err
return
}
}
}()
// control plane -> local service
go func() {
for {
msg, err := stream.Recv()
if err != nil {
errCh <- err
return
}
if msg.Close != nil {
errCh <- fmt.Errorf("server closed relay: %s", msg.Close.Reason)
return
}
if len(msg.Data) > 0 {
if _, err := conn.Write(msg.Data); err != nil {
errCh <- err
return
}
}
}
}()
err := <-errCh
_ = conn.Close()
_ = stream.CloseSend()
if errors.Is(err, io.EOF) {
return nil
}
return err
}
+72 -15
View File
@@ -24,6 +24,7 @@ import (
"gitea.hostxtra.co.uk/mrhid6/vantage/agent/internal/inventory"
"gitea.hostxtra.co.uk/mrhid6/vantage/agent/internal/keys"
"gitea.hostxtra.co.uk/mrhid6/vantage/agent/internal/monitors"
agentproxy "gitea.hostxtra.co.uk/mrhid6/vantage/agent/internal/proxy"
"gitea.hostxtra.co.uk/mrhid6/vantage/agent/internal/updates"
)
@@ -117,9 +118,21 @@ func poll(client *grpcclient.Client, cfg *config.Config, version string) error {
return nil
}
// How long a command stream must survive before it counts as having worked.
// Past this, the next drop is treated as a fresh incident rather than as the
// continuation of a run of failures.
const streamHealthyAfter = time.Minute
func runCommandStream(ctx context.Context, cfg *config.Config) {
backoff := time.Second
const maxBackoff = 2 * time.Minute
// Two minutes was the old ceiling, and it was reached far too easily. The
// command stream is what makes this agent controllable at all: while it is
// down, workflows and console sessions fail as "agent offline" even though
// SyncKeys keeps polling happily and the fleet list still shows the server
// active. A shorter ceiling costs a few reconnect attempts; the old one cost
// two minutes of an agent that looks fine and answers nothing.
const maxBackoff = 30 * time.Second
for {
select {
@@ -128,22 +141,40 @@ func runCommandStream(ctx context.Context, cfg *config.Config) {
default:
}
if err := connectAndHandleStream(ctx, cfg); err != nil {
if ctx.Err() != nil {
return
}
log.Printf("command stream error: %v, reconnecting in %s", err, backoff)
select {
case <-ctx.Done():
return
case <-time.After(backoff):
}
if backoff < maxBackoff {
backoff *= 2
}
} else {
started := time.Now()
err := connectAndHandleStream(ctx, cfg)
if ctx.Err() != nil {
return
}
// A stream that stayed up is evidence the control plane is reachable,
// whatever ended it. Without this the backoff only ever climbed:
// connectAndHandleStream returns an error on *every* stream end,
// including a healthy one dropped by a routine deploy, so an agent
// pinned itself at the ceiling after a handful of ordinary restarts and
// stayed there for the rest of its life.
if time.Since(started) >= streamHealthyAfter {
backoff = time.Second
}
if err != nil {
log.Printf("command stream error: %v, reconnecting in %s", err, backoff)
} else {
log.Printf("command stream closed, reconnecting in %s", backoff)
}
select {
case <-ctx.Done():
return
case <-time.After(backoff):
}
if backoff < maxBackoff {
backoff *= 2
if backoff > maxBackoff {
backoff = maxBackoff
}
}
}
}
@@ -197,6 +228,9 @@ func connectAndHandleStream(ctx context.Context, cfg *config.Config) error {
if cmd.CleanupWorkspace != nil {
go handleCleanupWorkspace(cmd)
}
if cmd.OpenProxy != nil {
go handleOpenProxy(ctx, cfg, cmd.OpenProxy)
}
if cmd.RunStep != nil {
go func(rc *pb.RunStepCmd, cid string) {
emit := func(seq uint64, data []byte) {
@@ -326,6 +360,29 @@ func handleCleanupWorkspace(cmd *pb.ServerCommand) {
log.Printf("removed run workspace %s (cmd=%s)", dir, cmd.CommandId)
}
// handleOpenProxy relays one console connection. It uses its own gRPC
// connection so console traffic never shares a stream with commands, key sync
// or workflow output.
func handleOpenProxy(ctx context.Context, cfg *config.Config, cmd *pb.OpenProxyCmd) {
client, err := grpcclient.New(cfg.ServerURL, cfg.TLS)
if err != nil {
log.Printf("proxy %s: dial control plane: %v", cmd.ProxyId, err)
return
}
defer client.Close()
stream, err := client.ProxyStream(ctx)
if err != nil {
log.Printf("proxy %s: open stream: %v", cmd.ProxyId, err)
return
}
log.Printf("proxy %s: relaying 127.0.0.1:%d", cmd.ProxyId, cmd.Port)
if err := agentproxy.Open(stream, cfg.ServerID, cfg.AgentToken, cmd.ProxyId, cmd.Port); err != nil {
log.Printf("proxy %s: %v", cmd.ProxyId, err)
}
}
func handleDeleteKey(cmd *pb.ServerCommand) {
label := cmd.DeleteKey.Label
keyPath := fmt.Sprintf("/root/.ssh/vantage_%s", strings.ReplaceAll(label, " ", "_"))
+144 -8
View File
@@ -120,7 +120,7 @@ Upload a public key, assign it per server, revoke softly. The agent diffs desire
### Workflows
A library of reusable **steps** (bash or PowerShell scripts with declared inputs, outputs, and secret refs) composed into **workflows** targeting a set of servers. Running one snapshots the resolved steps into a `WorkflowRun`, then dispatches `RunStepCmd` over the agent command stream. Step stdout/stderr streams back as `StepOutputChunk` and is written to a log file on disk; the UI streams it live. Steps support `on_failure: stop|continue|retry`, per-run env passed between steps via `output_env`, and a per-run workspace directory the agent cleans up at the end.
A library of reusable **steps** (bash or PowerShell scripts with declared inputs, outputs, and secret refs) composed into **workflows** targeting a set of servers. Running one snapshots the resolved steps into a `WorkflowRun`, then dispatches `RunStepCmd` over the agent command stream. Step stdout/stderr streams back as `StepOutputChunk` and is written to MongoDB (`workflow_log_lines`, one document per line); the UI streams it live. Steps support `on_failure: stop|continue|retry`, per-run env passed between steps via `output_env`, and a per-run workspace directory the agent cleans up at the end.
Default steps are seeded per org at boot (`SeedDefaultSteps`) from `VANTAGE_DEFAULT_STEPS_DIR`, which `server/Dockerfile` bakes to `/opt/default-steps` from the repo's `default_steps/`. Deliberately **not** under `/data` — that is a bind mount, so the library would be editable from the host. Adding a step there means committing a file and rebuilding, which is why `default_steps/` is in the `server` rebuild trigger. **Steps with `source: "default"` are read-only**: `UpdateStep`/`DeleteStep` refuse with `ErrDefaultStep` (409), because seeding rewrites them on every boot, so an edit would silently revert and a delete would come back. `web/` mirrors this — the step modal opens read-only, Delete is hidden, and the designer's per-step script override is `readOnly` for a default library step — but as elsewhere, the API is the boundary and the UI is the courtesy. Seeding writes straight to the collection rather than through `UpdateStep`, so the guard does not lock out the seeder. Logs are swept by retention (`workflow_log_retention_days`; nil = 30 days, 0 = forever).
@@ -138,7 +138,98 @@ Key/value pairs grouped by name, encrypted at rest with AES-256-GCM. Consumed tw
### Browser console
`POST /api/console/connect` mints a one-time session token; `GET /api/console/tunnel` upgrades to a WebSocket and proxies to **guacd** (Apache Guacamole daemon) using `github.com/wwt/guac`. SSH connections authenticate with a stored private key; RDP/VNC credentials are encrypted, single-use, and consumed when the tunnel opens.
`POST /api/console/connect` mints a one-time session token; `GET /api/console/tunnel`
upgrades to a WebSocket and proxies to **guacd** using `github.com/wwt/guac`.
guacd never dials the managed server. The server binds a single-use ephemeral
listener, pushes `OpenProxyCmd` down the agent's command stream, and the agent
opens a `ProxyStream` and relays the connection from its own **`127.0.0.1`** —
the host is hardcoded agent-side, so the control plane can name only a port.
This is what makes the console work on Vantage Cloud, where the customer's
server is behind NAT on a private address. It also means the console now
**requires a live agent** on every deployment: `consoleConnect` answers 409
`agent_offline` rather than hanging.
**guacd's Service is headless on purpose.** The server resolves `GUACD_ADDR` to
build the allow-list of sources permitted to claim a relay listener; a ClusterIP
resolves to the Service's virtual address while guacd connects from its *pod*
IP, so every relay connection is rejected and every session dies with
`waiting for guacd: i/o timeout`. Compose is immune — there the name resolves to
the address that connects.
SSH connections authenticate with a stored private key; RDP/VNC credentials are
encrypted, single-use, and consumed when the tunnel opens. None of them reach
the agent — the session is negotiated end-to-end between guacd and the target
daemon, so the agent relays bytes it cannot read.
### Running more than one server replica
An agent's `CommandStream` terminates on exactly **one** server process. Every
piece of coordination below exists because of that single fact: with several
replicas, the process asked to do something to an agent is almost never the
process holding that agent's stream.
`server/internal/bus` is the Redis message bus that closes the gap. It adds no
infrastructure — Redis was already required for sessions — and it is **not
optional on a single-replica deployment**: dispatch takes the bus path always,
so the code running in production is the code running everywhere, rather than a
rare cross-pod branch that only fails under load.
| Concern | How it crosses replicas |
| ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Which pod owns an agent | `vantage:agent:<server_id>` holds the owner's node ID with a 30s TTL, renewed every 10s. `Dispatcher.IsConnected` is an `EXISTS` on it |
| Sending a command | published to `vantage:cmd:<server_id>`; the owner pod acks on `vantage:ack:<command_id>`. **Request/ack, not a queue** — a command whose owner died must fail loudly (503) rather than queue |
| Step results | the owner pod publishes to `vantage:res:<command_id>`; the pod driving the run subscribes **before** dispatching, or a fast agent answers into a channel nobody has joined |
| Step output | never crosses. The dispatch envelope carries the secret mask list, so the owner pod masks and writes lines itself — unmasked bytes stay off the bus |
| Console relay | **not routed to the owner pod at all.** A `ProxyStream` is its own HTTP/2 request and an L7 proxy balances requests, not connections, so it does not follow the command stream — the listener therefore cannot be bound in advance. Whichever pod receives the stream binds it and announces **its own** address on `vantage:proxyaddr:<proxy_id>`; `vantage:proxypending:<proxy_id>` (30s, consumed atomically) is what authorises the claim, and the failure reason comes back on `vantage:proxyend:<proxy_id>` |
| Background jobs | `bus.RunAsLeader` — one Redis lock named `housekeeping` |
**Workflow logs are in MongoDB** (`workflow_log_lines`, one document per line,
with a `workflow_log_seq` counter document per run/server). Two pods write the
same log concurrently — the run's pod emits markers, the agent's pod emits
output — so ordering only means anything if both draw sequence numbers from the
same counter. `StepRun.log_offset` is that sequence number now, not a byte
offset. Writes are batched (128 lines or 250ms) and capped: 8 KB per line,
200k lines per server-run, after which one final `[vantage] log truncated`
marker is written and the rest is dropped. Without that cap a `yes` in a step
is a database incident. **Nothing writes to `/data` any more**, which is why
`server.persistence` now defaults to off and `VANTAGE_WORKFLOW_LOG_DIR` is gone.
**Shutdown order is load-bearing.** `main` traps SIGTERM, stops gRPC
(`GracefulStop`, 10s cap) and only then drains HTTP. Each `CommandStream`
handler releases its agent's presence claim on return, so a killed process
leaves `vantage:agent:<server_id>` behind for the rest of its 30s TTL — during
which other replicas dispatch to a pod that has exited and the caller sees
`agent offline` for a perfectly healthy agent. Draining HTTP first would hold
those claims for the length of the drain, which is why gRPC goes first. The
chart's `server.terminationGracePeriodSeconds` (30s) must stay above the
10s + 10s the stop sequence needs, or the kubelet SIGKILLs mid-shutdown and the
handling buys nothing.
The agent side of the same failure: `runCommandStream` resets its backoff only
after a stream that survived `streamHealthyAfter`. `connectAndHandleStream`
returns an error on *every* stream end, healthy ones included, so without that
reset the backoff only ever climbed — an agent pinned itself at the ceiling
after a handful of ordinary deploys and stayed there. The ceiling is 30s, not
minutes, because while the stream is down the agent still polls `SyncKeys` and
still reads as `active` in the fleet list while answering no commands at all.
**The leader lock is not an optimisation.** N replicas each running the monitor
scheduler means each check fires N times, each incident notification reaches the
customer N times, and each hourly rollup is written N times; N reapers race to
purge the same Free instance. `monitorsched`, `StartReaper`, `StartLogSweeper`,
`StartAuditSweeper` and the offline sweep therefore all run inside one
`RunAsLeader("housekeeping", …)` — one role, one lock. Each takes a context
cancelled the instant leadership is lost, and must return when it is.
Redis rather than a Kubernetes `Lease` so Compose takes the identical path: one
implementation to reason about, not two.
Two deployment requirements come with `replicaCount > 1`: every replica must
share **one** Redis (a per-pod Redis partitions the bus and every agent looks
offline to two thirds of the fleet), and `POD_IP` must be set — the chart does
it from the downward API — because `PROXY_ADVERTISE_HOST` names the Service, and
a Service cannot address the one pod holding a console listener.
### Inventory and OS updates
@@ -315,6 +406,7 @@ Key-state polling stays on the 30s `SyncKeys` interval. Full message definitions
Unauthenticated:
```
GET /healthz /readyz # liveness / readiness probes
GET /install /install.ps1 # dynamic agent install scripts
GET /update /update.ps1
GET /auth/bootstrap-status
@@ -421,7 +513,7 @@ Paddle is merchant of record; `admin/internal/paddle` is a thin REST client (no
## MongoDB Collections
`servers` · `keys` · `assignments` · `orgs` · `users` · `org_oidc` · `settings` · `secrets` · `workflows` · `workflow_steps` · `workflow_runs` · `monitors` · `incidents` · `monitor_rollups` · `notification_channels` · `console_sessions` · `audit_logs` · `migrations`
`servers` · `keys` · `assignments` · `orgs` · `users` · `org_oidc` · `settings` · `secrets` · `workflows` · `workflow_steps` · `workflow_runs` · `workflow_log_lines` · `workflow_log_seq` · `monitors` · `incidents` · `monitor_rollups` · `notification_channels` · `console_sessions` · `audit_logs` · `migrations`
Every document except `migrations` carries `org_id`. Struct definitions are the source of truth — see `server/internal/models/`.
@@ -433,6 +525,7 @@ Notes that are not obvious from the structs:
- `assignments.revoked_at: null` means active. Revocation is soft, preserving audit history.
- `workflow_runs.steps_snapshot` freezes the resolved steps so editing the library never rewrites history.
- `console_sessions.token_consumed_at` is set atomically to enforce one-time use.
- `workflow_log_lines` is keyed `(run_id, server_id, seq)` — the index is not an optimisation, every read is a range scan over it. `workflow_log_seq` holds one counter document per `run_id/server_id`, which is what lets two pods interleave into one ordered log. Neither carries `instance_id`: they are reached only through a run, and a run is already scoped.
- `users.auth_source` is `local`, `oidc` or `hq`. An `hq` user was projected from a Vantage HQ account and carries `hq_user_id`; HQ owns its role, password and existence.
Admin's own database is separate and holds `accounts` · `admin_instances` · `licenses` · `subscriptions` · `plans` · `catalogue` · `entitlements` · `paddle_events` · `staff_users` · `customer_users` · `instance_members` · `admin_audit`. `paddle_events` is the webhook idempotency log, unique on `event_id`: an event is claimed there before processing, and a duplicate of a handled event is a 200 no-op. `instance_members` is unique on `(instance_id, customer_user_id)` — one person holds at most one user in one instance, which makes a grant idempotent-by-refusal rather than silently doubling a projection. It is an _index_ of the control-plane rows, not the authority (see "Grants project, they do not federate"). Admin has no migrations collection; `models.Backfill` runs on every boot and is idempotent by filtering on the absence of what it writes.
@@ -514,11 +607,17 @@ Windows: MSI built by CI (WiX), or `installer/setup.ps1` registering the agent a
| `GRPC_HOST` | **yes** | `host:port` agents dial. Boot fails without it — there is no safe default; falling back to the web host would hand agents a port that does not speak gRPC. |
| `MONGO_URI` | no | default `mongodb://localhost:27017` |
| `MONGO_DB` | no | default `vantage` |
| `REDIS_USERNAME` | no | Redis 6+ ACL user. Leave empty for a legacy `requirepass` instance — go-redis then sends AUTH with one argument instead of two |
| `REDIS_PASSWORD` | no | empty for an unauthenticated Redis |
| `REDIS_ADDR` | no | default `localhost:6379` |
| `KEY_ENCRYPTION_KEY` | yes in practice | 64-char hex (32 bytes) for AES-256-GCM. Required for private keys, secrets, OIDC secrets, RDP credentials. |
| `GUACD_ADDR` | no | default `guacd:4822` |
| `PROXY_ADVERTISE_HOST` | no | default `server`; the hostname guacd resolves the control plane by, handed to guacd as the relay's address. Wrong here and every console session fails at connect |
| `PROXY_LISTEN_HOST` | no | default `0.0.0.0`; the interface the ephemeral relay listener binds |
| `APP_ROOT_LABEL` | no | default `vantage`; wrong value disables the host/session org guard |
| `VANTAGE_WORKFLOW_LOG_DIR` | no | where run logs are written |
| `POD_IP` | no | this pod's own address, set by the Helm chart from the downward API. **Takes precedence over `PROXY_ADVERTISE_HOST`** — a console relay listener belongs to one replica, and a Service address names all of them |
| `VANTAGE_MIGRATE_ONLY` | no | run schema setup (migrations, index builders, default-step seeding) and exit without serving. `GRPC_HOST` is not required in this mode. Set by the Helm chart's pre-upgrade Job |
| `VANTAGE_SKIP_MIGRATIONS` | no | serve without running schema setup, on the assumption a Job already did. Set by the chart's Deployment whenever `server.migrationJob.enabled`. Unset under Compose, where one process still migrates and then serves |
| `FREE_INSTANCE_REAP_AFTER` | no | duration past a Free licence's expiry before the instance and all its data are deleted. **Empty disables the reaper, and empty is the default.** Set to `336h` in `docker-compose.site.yml` only — a self-hosted deployment must never reap. Must match admin's value, which only names the date in warning emails |
**sitesvc** (`deploy/docker-compose.site.yml` only):
@@ -533,11 +632,33 @@ Windows: MSI built by CI (WiX), or `installer/setup.ps1` registering the agent a
| `SITE_ORIGIN` | yes in practice | comma-separated allowed origins; unset refuses every cross-origin browser request |
| `TRUST_PROXY` | no | only `true` behind a proxy that overwrites `X-Forwarded-For`, or clients spoof past the rate limiter |
### Ingress (Helm, Traefik)
`ingress.enabled` publishes **two** hostnames, because the two audiences arrive over different protocols:
| Values | Route |
| -------------------------------------------- | --------------------------------------------------------------------------------------------- |
| `ingress.web.host` (+ `web.extraHosts`) | browsers → `web:3000` |
| `ingress.api.paths` (when `api.enabled`) | `/api`, `/auth` → `<release>-server:8080`, bypassing the Next proxy |
| `ingress.grpc.host` | agents → a dedicated `<release>-server-grpc` Service on 9090, annotated `serversscheme: h2c` |
**`ingress.web.host` is normally a wildcard.** `*.vantage.example.com` is the per-tenant instance namespace — `APP_ROOT_LABEL` resolves the instance from the label. A Kubernetes wildcard host matches **exactly one** label, so it does not match the apex, and here that is correct rather than a gap: `vantage.hostxtra.co.uk` is the marketing site (`site/`, in `docker-compose.site.yml`), which this chart does not deploy. `extraHosts` is for a genuine second name; adding the apex to it would put the control plane on the marketing host. Every host in the list gets identical paths.
**`ingress.api.enabled` routes `/api` and `/auth` straight to the server.** Both arrangements work — without it `web` proxies those prefixes onward itself (`web/next.config.ts`) — but edge routing is one hop shorter and matches what the Nginx Proxy Manager in front of the Docker deployment already does, so leaving it off makes the request path a different shape on Kubernetes than in production. It stays **off by default** because it only helps where the server is reachable on the same host and certificate as `web`; turning it on blindly moves the whole API onto a route that may not be provisioned. Traefik derives router priority from rule length, so `PathPrefix(/api)` outranks the catch-all `/` with no priority annotation needed.
**The gRPC route needs its own Service.** The server terminates no TLS; it speaks plain h2c and always has, with TLS terminated by whatever sits in front. Traefik will not use h2c to a backend unless the *Service* says so, and that annotation applies to every port on the Service — so annotating the shared two-port `<release>-server` would force h2c on its HTTP port too.
**`server.env.grpcHost` is not derived from `ingress.grpc.host`, and the chart refuses to render if they disagree.** Agents dial whatever `grpcHost` says, and it is baked into every install one-liner; left pointing at the in-cluster Service while agents arrive through the ingress, every install succeeds and every agent then fails to connect, with nothing in the control plane explaining why. Guessing at the port (443? 9090?) would be worse than stopping.
TLS is `ingress.tls.secretName` / `grpcSecretName` (pre-existing certificates) **or** `certResolver` (Traefik ACME). Setting neither while `tls.enabled` produces a TLS router with no certificate, so Traefik serves its self-signed default — which looks valid and is trusted by nothing. NOTES.txt warns on install rather than the chart failing, since it is a real if unusual choice behind another terminator.
---
`deploy/docker-compose.yml` runs four services: `redis`, `guacd`, `server` (8080 + 9090), `web` (3000). MongoDB is external. `deploy/docker-compose.site.yml` adds five more — `site` (3003), `sitesvc` (8082), `admin` (8083), `adminsite` (3004) and `docsite` (3005) — and is only used on vantage.hostxtra.co.uk.
`docsite` is the odd one: a **static** build served by `nginx:alpine-slim`, not a Node runtime, and it listens on `80` rather than `3000`. It is reached at **`vantage.hostxtra.co.uk/docs`** — a path on the marketing host, routed by its own Nginx Proxy Manager location, which must sort **above** the catch-all forwarding to `site:3003` or Next answers the 404. A path and not a subdomain because `*.vantage.hostxtra.co.uk` is the per-tenant instance namespace and `APP_ROOT_LABEL` would read a `docs.` label as a tenant slug. NPM forwards the **full** path upstream — it does not strip `/docs` — so `DOCS_BASE_URL`, the proxy location and the directory the image copies the build into (`/usr/share/nginx/html/docs`) must all agree. When they do not, the HTML loads and every asset 404s.
`LICENSE_SIGNING_KEY` appears in **exactly one service in exactly one compose file**: `admin` in `docker-compose.site.yml`. It must never be added to `server`, and the self-hosted `docker-compose.yml` must never mention `admin` or `adminsite` at all. Admin uses an external Redis via `REDIS_ADDR`/`REDIS_USERNAME`/`REDIS_PASSWORD`; the base compose hardcodes `redis:6379` for `server`, so those variables reach admin only.
`LICENSE_SIGNING_KEY` appears in **exactly one service in exactly one compose file**: `admin` in `docker-compose.site.yml`. It must never be added to `server`, and the self-hosted `docker-compose.yml` must never mention `admin` or `adminsite` at all. Admin uses an external Redis via `REDIS_ADDR`/`REDIS_USERNAME`/`REDIS_PASSWORD`. `server` now reads the same three, so a Kubernetes install can point at a managed Redis; the base compose still hardcodes an unauthenticated `redis:6379` for it, so in Docker those credentials remain admin's alone.
---
@@ -651,10 +772,24 @@ cd /opt/vantage && docker compose -f docker-compose.yml -f docker-compose.site.y
The gap this leaves: **changing a repo variable pushes no commit, so nothing rebuilds.** After editing `ADMIN_API_URL`, `HQ_URL` or `ADMIN_ENV`, run the workflow manually — that is what `workflow_dispatch` is there for. Base images also stop being refreshed on a service nobody touches; a periodic manual run covers that.
### `chart-release.yml` — validates on every chart change, publishes on `chart/v*` tags
Two jobs' worth of work in one, split by trigger. Any push or PR touching `deploy/chart/` lints the chart and renders it four ways: defaults, a multi-replica install, external Redis and MongoDB, and a set of values that **must be refused**. That last one is the point — every safety rail in this chart is a template `fail`, and `helm lint` happily accepts a chart whose templates never execute, so only rendering proves they still fire.
Publishing runs only on a `chart/v*` tag, to the Gitea Helm registry at `/api/packages/<owner>/helm/api/charts`. **`Chart.yaml` is the source of truth for the version**; the tag only selects which one to publish, and a tag that disagrees with `Chart.yaml` fails rather than stamping over it — the alternative leaves the repository disagreeing with what shipped. A version already in the registry is rejected by Gitea, which is intended: published chart versions are immutable.
The registry host comes from `github.server_url`, so it cannot drift from the instance the workflow is running on. It authenticates with `REGISTRY_USER` + **`RELEASE_TOKEN`** — the pair `server-deploy.yml` actually uses for `docker login`. `REGISTRY_PASSWORD` is listed in the secrets table below but set by no workflow; passing an unset secret yields an empty password and Gitea answers `401 Failed to authenticate user`, which reads like a scope problem on a token that was never sent. The publish step therefore checks both are non-empty before it calls curl. `RELEASE_TOKEN` needs `write:package` in addition to `write:release`.
```bash
helm repo add vantage https://gitea.hostxtra.co.uk/api/packages/mrhid6/helm
helm install vantage vantage/vantage --version 0.1.0
```
### Tagging
```bash
git tag agent/v1.0.0 && git push origin agent/v1.0.0 # agent release
git tag chart/v0.1.0 && git push origin chart/v0.1.0 # helm chart package
git push origin main # server + web deploy
```
@@ -662,9 +797,9 @@ git push origin main # server + web deploy
| Name | Type | Value |
| ----------------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `RELEASE_TOKEN` | Secret | Gitea API token, `write:release` |
| `REGISTRY_USER` | Secret | Gitea username |
| `REGISTRY_PASSWORD` | Secret | Gitea token, `write:packages` |
| `RELEASE_TOKEN` | Secret | Gitea API token. Needs `write:release` (agent releases), `write:package` (container images and the Helm chart). **This is the only token any workflow authenticates with** — `docker login` and the chart publish both pair it with `REGISTRY_USER` |
| `REGISTRY_USER` | Secret | Gitea username. Must own `RELEASE_TOKEN`, or basic auth is rejected |
| ~~`REGISTRY_PASSWORD`~~ | — | **Not used.** Named here historically; no workflow reads it. Referencing an unset secret yields an empty password and a `401 Failed to authenticate user` that looks like a token scope problem. Use `RELEASE_TOKEN` |
| `DOCKER_HOST` | Variable | registry host used for image tags |
| `API_URL` | **not** a CI variable | `web` reads it at **runtime**, from the container environment — `next.config.ts` is evaluated when `server.js` boots in standalone mode, and the rewrites it feeds are server-side, never browser-side. Default `http://localhost:8080`; compose sets `http://server:8080`. `NEXT_PUBLIC_API_URL` is still honoured as a fallback for existing deployments. |
| `SITE_API_URL` | Variable | **browser-reachable** sitesvc URL, baked into the `site` image. Required — if empty, both forms report "not connected" and submit nowhere. Must also be in sitesvc's `SITE_ORIGIN`. |
@@ -698,6 +833,7 @@ git push origin main # server + web deploy
- **`org_id` on every document** — isolation enforced at the query layer, not by separate databases.
- **root only** — manages `/root/.ssh/authorized_keys`; no per-user key management.
- **Windows agents are second-class by design** — register, heartbeat, run steps, report inventory; no `authorized_keys` management.
- **Both `server` and `web` scale horizontally** — see "Running more than one server replica" below. `web` holds nothing; `server` holds per-agent state that is routed between replicas over Redis rather than duplicated.
- **Deletion lives in the control plane** — admin sends the warnings because it knows the billing address; the control plane performs the delete because it is the only service that knows which collections carry `instance_id`. Mirroring that list into admin would drift, and a drift there deletes the wrong rows.
## graphify
-23
View File
@@ -1,23 +0,0 @@
[Unit]
Description=Vantage Agent
Documentation=https://github.com/your-org/vantage
After=network.target
Wants=network-online.target
[Service]
Type=simple
ExecStart=/usr/local/bin/vantage-agent
Restart=always
RestartSec=10
User=root
StandardOutput=journal
StandardError=journal
SyslogIdentifier=vantage-agent
# Security hardening
NoNewPrivileges=true
ProtectSystem=false
ProtectHome=false
[Install]
WantedBy=multi-user.target
+6
View File
@@ -0,0 +1,6 @@
apiVersion: v2
name: vantage
description: Helm chart for the Vantage stack (Redis, MongoDB, guacd, server, web)
type: application
version: 1.0.7
appVersion: "1.0.7"
+72
View File
@@ -0,0 +1,72 @@
Vantage has been deployed as release "{{ .Release.Name }}" in namespace "{{ .Release.Namespace }}".
Services created:
{{- if .Values.redis.enabled }}
- {{ .Release.Name }}-redis (ClusterIP {{ .Values.redis.port }})
{{- else }}
- Redis: not deployed, using external {{ .Values.redis.addr }}
{{- end }}
{{- if .Values.mongo.enabled }}
- {{ .Release.Name }}-mongo (ClusterIP {{ .Values.mongo.port }})
{{- else }}
- MongoDB: not deployed, using the external server.env.mongoUri
{{- end }}
- {{ .Release.Name }}-guacd ({{ .Values.guacd.service.type }} {{ .Values.guacd.service.port }})
- {{ .Release.Name }}-server ({{ .Values.server.service.type }} http:{{ .Values.server.service.httpPort }} grpc:{{ .Values.server.service.grpcPort }})
- {{ .Release.Name }}-web ({{ .Values.web.service.type }} {{ .Values.web.service.port }})
Scaling (server.replicaCount / web.replicaCount):
- Both scale. Pin the image tags first — replicas on different builds serve
mismatched web asset hashes, and mixed server versions share one bus.
- server replicas route agent commands, step results and console relays to
each other over Redis, so every replica must use the SAME Redis. Workflow
logs are in MongoDB, not on a volume.
- Background work (monitor scheduler, Free reaper, log and audit retention,
the offline sweep) runs on one replica at a time under a Redis leader lock.
- server.persistence must be off to scale past one replica on a ReadWriteOnce
volume. Nothing writes to it any more.
{{- if gt (int .Values.server.replicaCount) 1 }}
- Console relays are reached by pod IP; guacd must be able to dial pod IPs
directly (it can, inside the cluster network).
{{- end }}
{{- if .Values.server.migrationJob.enabled }}
- Migrations run in the {{ .Release.Name }}-migrate Job before each upgrade;
the pods skip them. Its logs are kept: kubectl logs job/{{ .Release.Name }}-migrate
{{- end }}
{{- if .Values.ingress.enabled }}
Ingress (Traefik):
- Browsers: https://{{ .Values.ingress.web.host }}
{{- range .Values.ingress.web.extraHosts }}
https://{{ . }}
{{- end }}
{{- if .Values.ingress.api.enabled }}
{{ join ", " .Values.ingress.api.paths }} go straight to the server; everything else to web.
{{- else }}
Everything goes to web, which proxies /api and /auth onward. Set
ingress.api.enabled=true to route them at the edge instead.
{{- end }}
{{- if .Values.ingress.grpc.enabled }}
- Agents: {{ .Values.ingress.grpc.host }} (gRPC, h2c behind TLS)
Agents dial server.env.grpcHost, currently {{ tpl .Values.server.env.grpcHost . }}.
Point DNS for both hostnames at the Traefik load balancer.
{{- if not .Values.ingress.tls.enabled }}
- WARNING: ingress.tls.enabled is false. Agent tokens and session cookies
would cross the network in clear.
{{- else if and (not .Values.ingress.tls.certResolver) (not .Values.ingress.tls.secretName) }}
- WARNING: TLS is on but neither ingress.tls.secretName nor
ingress.tls.certResolver is set, so Traefik will serve its self-signed
default certificate, which no agent and no browser will trust.
{{- end }}
{{- end }}
{{- end }}
By default the server/web/guacd services are ClusterIP only (no host port publishing,
unlike the original docker-compose file). To expose them externally, set
server.service.type / web.service.type / guacd.service.type to NodePort or LoadBalancer,
or add an Ingress on top of the -web and -server services.
Quick access via port-forward, e.g.:
kubectl port-forward svc/{{ .Release.Name }}-web {{ .Values.web.service.port }}:{{ .Values.web.service.port }}
kubectl port-forward svc/{{ .Release.Name }}-server {{ .Values.server.service.httpPort }}:{{ .Values.server.service.httpPort }}
@@ -0,0 +1,85 @@
{{/*
Common name helpers
*/}}
{{- define "vantage.fullname" -}}
{{ .Release.Name }}
{{- end -}}
{{- define "vantage.labels" -}}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/managed-by: {{ .Release.Service }}
{{- end -}}
{{/*
vantage.server.env renders the server container's environment.
It lives here because two workloads need it identically: the Deployment and the
pre-upgrade migration Job. A Job that connected to a different database than the
pods it migrates for would be worse than no Job at all, so there is one copy and
both read it.
*/}}
{{- define "vantage.server.env" -}}
- name: MONGO_URI
{{- $mongoUri := tpl .Values.server.env.mongoUri . }}
{{- if and (not .Values.mongo.enabled) (contains (printf "%s-mongo" .Release.Name) $mongoUri) }}
{{- fail "mongo.enabled is false, so server.env.mongoUri must point at an external MongoDB rather than the in-chart one" }}
{{- end }}
value: {{ $mongoUri | quote }}
- name: REDIS_ADDR
{{- if .Values.redis.enabled }}
value: "{{ .Release.Name }}-redis:{{ .Values.redis.port }}"
{{- else }}
{{- if not .Values.redis.addr }}
{{- fail "redis.enabled is false, so redis.addr must be set to an external Redis host:port" }}
{{- end }}
value: {{ .Values.redis.addr | quote }}
{{- end }}
{{- if .Values.redis.auth.existingSecret }}
- name: REDIS_USERNAME
valueFrom:
secretKeyRef:
name: {{ .Values.redis.auth.existingSecret }}
key: {{ .Values.redis.auth.usernameKey }}
optional: true
- name: REDIS_PASSWORD
valueFrom:
secretKeyRef:
name: {{ .Values.redis.auth.existingSecret }}
key: {{ .Values.redis.auth.passwordKey }}
{{- else }}
{{- if .Values.redis.auth.username }}
- name: REDIS_USERNAME
value: {{ .Values.redis.auth.username | quote }}
{{- end }}
{{- if .Values.redis.auth.password }}
- name: REDIS_PASSWORD
value: {{ .Values.redis.auth.password | quote }}
{{- end }}
{{- end }}
- name: GRPC_HOST
value: {{ .Values.server.env.grpcHost | quote }}
- name: GRPC_PORT
value: {{ .Values.server.service.grpcPort | quote }}
- name: HTTP_PORT
value: {{ .Values.server.service.httpPort | quote }}
- name: KEY_ENCRYPTION_KEY
value: {{ .Values.server.env.keyEncryptionKey | quote }}
- name: GUACD_ADDR
value: "{{ .Release.Name }}-guacd:{{ .Values.guacd.service.port }}"
- name: APP_ROOT_LABEL
value: {{ .Values.server.env.appRootLabel | quote }}
- name: PROXY_ADVERTISE_HOST
value: {{ .Values.server.env.proxyAdvertiseHost | quote }}
- name: PROXY_LISTEN_HOST
value: {{ .Values.server.env.proxyListenHost | quote }}
{{- if eq .Values.server.env.deploymentType "cloud" }}
- name: VANTAGE_DEPLOYMENT
value: "cloud"
- name: FREE_INSTANCE_REAP_AFTER
value: {{ .Values.server.env.freeInstanceReapAfter | quote }}
{{- end }}
- name: POD_IP
valueFrom:
fieldRef:
fieldPath: status.podIP
{{- end -}}
+47
View File
@@ -0,0 +1,47 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: {{ .Release.Name }}-guacd
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: guacd
spec:
replicas: 1
selector:
matchLabels:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: guacd
template:
metadata:
labels:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: guacd
spec:
{{- if .Values.imagePullSecrets }}
imagePullSecrets:
{{- toYaml .Values.imagePullSecrets | nindent 8 }}
{{- end }}
containers:
- name: guacd
image: "{{ .Values.guacd.image.repository }}:{{ .Values.guacd.image.tag }}"
ports:
- containerPort: 4822
---
apiVersion: v1
kind: Service
metadata:
name: {{ .Release.Name }}-guacd
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: guacd
spec:
type: {{ .Values.guacd.service.type }}
{{- if eq .Values.guacd.service.type "ClusterIP" }}
clusterIP: None
{{- end }}
selector:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: guacd
ports:
- port: {{ .Values.guacd.service.port }}
targetPort: 4822
+173
View File
@@ -0,0 +1,173 @@
{{- if .Values.ingress.enabled }}
{{/*
Two hostnames, because the two audiences arrive over different protocols.
Browsers reach the web host. What answers there depends on the path: with
ingress.api.enabled, /api and /auth go straight to the server and everything
else goes to `web`. Without it, everything goes to `web`, which proxies those
prefixes onward itself (web/next.config.ts).
Both work. Routing at the edge is one hop shorter and is what the Nginx Proxy
Manager deployment in front of the Docker install already does, so leaving it
off changes the shape of the request path between the two deployments. It is
still off by default, because turning it on where `web` is the only thing with
a public certificate would strand /api behind a route nobody can reach.
The web host is normally a wildcard — `*.vantage.example.com` — because that is
the per-tenant instance namespace; APP_ROOT_LABEL resolves the instance from the
label. Kubernetes wildcard hosts match exactly one label, so this does not match
the apex, and on the deployment this chart was written for it must not: the apex
is the marketing site, a separate application (see site/ and
docker-compose.site.yml). extraHosts exists for a genuine second name, not for
reclaiming the apex.
Agents reach the server's gRPC port, which is plain h2c — the server holds no
certificates of its own, TLS has always been terminated by whatever sits in
front. Traefik will not speak h2c to a backend unless told to, and it is told
per Service, which is why the gRPC route gets a Service of its own below rather
than reusing the two-port one. Annotating the shared Service would force h2c on
its HTTP port too.
*/}}
{{- $tls := .Values.ingress.tls }}
{{- $webHost := required "ingress.enabled requires ingress.web.host" .Values.ingress.web.host }}
{{- $hosts := prepend .Values.ingress.web.extraHosts $webHost }}
{{- $apiPaths := .Values.ingress.api.paths }}
{{- if and .Values.ingress.api.enabled (not $apiPaths) }}
{{- fail "ingress.api.enabled requires at least one path in ingress.api.paths" }}
{{- end }}
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: {{ .Release.Name }}-web
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: web
annotations:
traefik.ingress.kubernetes.io/router.entrypoints: {{ .Values.ingress.entrypoint | quote }}
{{- if and $tls.enabled $tls.certResolver }}
traefik.ingress.kubernetes.io/router.tls: "true"
traefik.ingress.kubernetes.io/router.tls.certresolver: {{ $tls.certResolver | quote }}
{{- else if $tls.enabled }}
traefik.ingress.kubernetes.io/router.tls: "true"
{{- end }}
{{- with .Values.ingress.annotations }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
{{- if .Values.ingress.className }}
ingressClassName: {{ .Values.ingress.className }}
{{- end }}
{{- if and $tls.enabled $tls.secretName }}
tls:
- hosts:
{{- range $hosts }}
- {{ . | quote }}
{{- end }}
secretName: {{ $tls.secretName }}
{{- end }}
rules:
{{- range $host := $hosts }}
- host: {{ $host | quote }}
http:
paths:
{{- /*
The API paths come first and, more importantly, are longer. Traefik
derives router priority from rule length, so Host(x) &&
PathPrefix(/api) outranks Host(x) && PathPrefix(/) without anyone
having to assign priorities by hand. Order within the list is for
the reader; the length is what decides.
*/}}
{{- if $.Values.ingress.api.enabled }}
{{- range $apiPaths }}
- path: {{ . | quote }}
pathType: Prefix
backend:
service:
name: {{ $.Release.Name }}-server
port:
number: {{ $.Values.server.service.httpPort }}
{{- end }}
{{- end }}
- path: /
pathType: Prefix
backend:
service:
name: {{ $.Release.Name }}-web
port:
number: {{ $.Values.web.service.port }}
{{- end }}
{{- if .Values.ingress.grpc.enabled }}
{{- $grpcHost := required "ingress.grpc.enabled requires ingress.grpc.host" .Values.ingress.grpc.host }}
{{/*
GRPC_HOST is what an agent is told to dial, and it is baked into every install
one-liner. Left pointing at the in-cluster Service while agents are expected to
arrive through the ingress, every install would succeed and every agent would
fail to connect — with nothing in the control plane saying why.
*/}}
{{- $grpcEnv := tpl .Values.server.env.grpcHost . }}
{{- if contains (printf "%s-server" .Release.Name) $grpcEnv }}
{{- fail (printf "ingress.grpc.enabled routes agents through %s, but server.env.grpcHost is still the in-cluster address %q. Agents dial the value of grpcHost, so set it to the public gRPC address (for example %q)." $grpcHost $grpcEnv (printf "%s:443" $grpcHost)) }}
{{- end }}
---
apiVersion: v1
kind: Service
metadata:
name: {{ .Release.Name }}-server-grpc
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: server
annotations:
# The server speaks h2c: it terminates no TLS itself. Without this Traefik
# dials the backend as HTTP/1.1 and every agent handshake fails.
traefik.ingress.kubernetes.io/service.serversscheme: h2c
spec:
type: ClusterIP
selector:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: server
ports:
- name: grpc
port: {{ .Values.server.service.grpcPort }}
targetPort: {{ .Values.server.service.grpcPort }}
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: {{ .Release.Name }}-grpc
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: server
annotations:
traefik.ingress.kubernetes.io/router.entrypoints: {{ .Values.ingress.entrypoint | quote }}
{{- if and $tls.enabled $tls.certResolver }}
traefik.ingress.kubernetes.io/router.tls: "true"
traefik.ingress.kubernetes.io/router.tls.certresolver: {{ $tls.certResolver | quote }}
{{- else if $tls.enabled }}
traefik.ingress.kubernetes.io/router.tls: "true"
{{- end }}
{{- with .Values.ingress.grpc.annotations }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
{{- if .Values.ingress.className }}
ingressClassName: {{ .Values.ingress.className }}
{{- end }}
{{- if and $tls.enabled $tls.grpcSecretName }}
tls:
- hosts:
- {{ $grpcHost | quote }}
secretName: {{ $tls.grpcSecretName }}
{{- end }}
rules:
- host: {{ $grpcHost | quote }}
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: {{ .Release.Name }}-server-grpc
port:
number: {{ .Values.server.service.grpcPort }}
{{- end }}
{{- end }}
@@ -0,0 +1,68 @@
{{- if .Values.server.migrationJob.enabled }}
{{/*
Schema setup, lifted out of the serving pods.
Every server process used to run migrations, index builders and default-step
seeding at boot. With one replica that is fine. With two it is not: 0004 renames
the orgs collection to instances, and a sibling reading it mid-rename is a
corruption, not a retry.
A Helm hook Job runs it once, before any pod of the new version starts. The
Deployment then sets VANTAGE_SKIP_MIGRATIONS, which is what makes the Job's
existence load-bearing rather than decorative — if you disable the Job, the
pods go back to migrating themselves and you must go back to one replica.
hook-weight orders this after the dependency waits; before-hook-creation deletes
the previous Job so a repeat upgrade is not blocked by an immutable object. The
Job is deliberately NOT deleted on success: its logs are the record of what the
upgrade did to the database.
*/}}
apiVersion: batch/v1
kind: Job
metadata:
name: {{ .Release.Name }}-migrate
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: migrate
annotations:
"helm.sh/hook": pre-install,pre-upgrade
"helm.sh/hook-weight": "0"
"helm.sh/hook-delete-policy": before-hook-creation
spec:
backoffLimit: {{ .Values.server.migrationJob.backoffLimit }}
# A migration that has not finished in this long is stuck, and a stuck
# migration should fail the upgrade rather than hold it open forever.
activeDeadlineSeconds: {{ .Values.server.migrationJob.activeDeadlineSeconds }}
template:
metadata:
labels:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: migrate
spec:
restartPolicy: Never
{{- if .Values.imagePullSecrets }}
imagePullSecrets:
{{- toYaml .Values.imagePullSecrets | nindent 8 }}
{{- end }}
{{- if .Values.mongo.enabled }}
# Only Mongo. The Job never opens Redis, and waiting on a Redis this
# chart may not even deploy would block an upgrade for no reason.
initContainers:
- name: wait-for-mongo
image: busybox:1.36
command:
- sh
- -c
- |
until nc -z {{ .Release.Name }}-mongo {{ .Values.mongo.port }}; do
echo "waiting for mongo..."; sleep 2;
done
{{- end }}
containers:
- name: migrate
image: "{{ .Values.server.image.repository }}:{{ .Values.server.image.tag }}"
env:
{{- include "vantage.server.env" . | nindent 12 }}
- name: VANTAGE_MIGRATE_ONLY
value: "true"
{{- end }}
+92
View File
@@ -0,0 +1,92 @@
{{- if .Values.mongo.enabled }}
{{- if .Values.mongo.persistence.enabled }}
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: {{ .Release.Name }}-mongo-data
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: mongo
spec:
accessModes:
- {{ .Values.mongo.persistence.accessMode }}
{{- if .Values.mongo.persistence.storageClass }}
storageClassName: {{ .Values.mongo.persistence.storageClass }}
{{- end }}
resources:
requests:
storage: {{ .Values.mongo.persistence.size }}
---
{{- end }}
apiVersion: apps/v1
kind: Deployment
metadata:
name: {{ .Release.Name }}-mongo
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: mongo
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: mongo
template:
metadata:
labels:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: mongo
spec:
{{- if .Values.imagePullSecrets }}
imagePullSecrets:
{{- toYaml .Values.imagePullSecrets | nindent 8 }}
{{- end }}
containers:
- name: mongo
image: "{{ .Values.mongo.image.repository }}:{{ .Values.mongo.image.tag }}"
ports:
- containerPort: {{ .Values.mongo.port }}
volumeMounts:
- name: mongo-data
mountPath: /data/db
livenessProbe:
exec:
command: ["mongosh", "--quiet", "--eval", "db.adminCommand('ping')"]
initialDelaySeconds: 10
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 5
readinessProbe:
exec:
command: ["mongosh", "--quiet", "--eval", "db.adminCommand('ping')"]
initialDelaySeconds: 5
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 5
volumes:
- name: mongo-data
{{- if .Values.mongo.persistence.enabled }}
persistentVolumeClaim:
claimName: {{ .Release.Name }}-mongo-data
{{- else }}
emptyDir: {}
{{- end }}
---
apiVersion: v1
kind: Service
metadata:
name: {{ .Release.Name }}-mongo
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: mongo
spec:
type: ClusterIP
selector:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: mongo
ports:
- port: {{ .Values.mongo.port }}
targetPort: {{ .Values.mongo.port }}
{{- end }}
+92
View File
@@ -0,0 +1,92 @@
{{- if .Values.redis.enabled }}
{{- if .Values.redis.persistence.enabled }}
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: {{ .Release.Name }}-redis-data
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: redis
spec:
accessModes:
- {{ .Values.redis.persistence.accessMode }}
{{- if .Values.redis.persistence.storageClass }}
storageClassName: {{ .Values.redis.persistence.storageClass }}
{{- end }}
resources:
requests:
storage: {{ .Values.redis.persistence.size }}
---
{{- end }}
apiVersion: apps/v1
kind: Deployment
metadata:
name: {{ .Release.Name }}-redis
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: redis
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: redis
template:
metadata:
labels:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: redis
spec:
{{- if .Values.imagePullSecrets }}
imagePullSecrets:
{{- toYaml .Values.imagePullSecrets | nindent 8 }}
{{- end }}
containers:
- name: redis
image: "{{ .Values.redis.image.repository }}:{{ .Values.redis.image.tag }}"
ports:
- containerPort: {{ .Values.redis.port }}
volumeMounts:
- name: redis-data
mountPath: /data
livenessProbe:
exec:
command: ["redis-cli", "ping"]
initialDelaySeconds: 10
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 5
readinessProbe:
exec:
command: ["redis-cli", "ping"]
initialDelaySeconds: 5
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 5
volumes:
- name: redis-data
{{- if .Values.redis.persistence.enabled }}
persistentVolumeClaim:
claimName: {{ .Release.Name }}-redis-data
{{- else }}
emptyDir: {}
{{- end }}
---
apiVersion: v1
kind: Service
metadata:
name: {{ .Release.Name }}-redis
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: redis
spec:
type: ClusterIP
selector:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: redis
ports:
- port: {{ .Values.redis.port }}
targetPort: {{ .Values.redis.port }}
{{- end }}
+181
View File
@@ -0,0 +1,181 @@
{{- if and .Values.server.persistence.enabled (not .Values.server.persistence.useHostPath) }}
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: {{ .Release.Name }}-server-data
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: server
spec:
accessModes:
- {{ .Values.server.persistence.accessMode }}
{{- if .Values.server.persistence.storageClass }}
storageClassName: {{ .Values.server.persistence.storageClass }}
{{- end }}
resources:
requests:
storage: {{ .Values.server.persistence.size }}
---
{{- end }}
apiVersion: apps/v1
kind: Deployment
metadata:
name: {{ .Release.Name }}-server
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: server
spec:
{{- $replicas := int .Values.server.replicaCount }}
replicas: {{ $replicas }}
{{- if and .Values.server.persistence.enabled (eq .Values.server.persistence.accessMode "ReadWriteOnce") }}
# A ReadWriteOnce volume cannot be mounted by a second pod at all, and cannot
# be handed to a new pod while the old one still holds it. Persistence is off
# by default now that nothing writes to it; if it is on, replicas are capped
# at one and updates go through Recreate.
{{- if gt $replicas 1 }}
{{- fail "server.persistence.enabled with a ReadWriteOnce volume cannot be combined with server.replicaCount > 1. Nothing in the server writes to that volume any more (workflow logs live in MongoDB); set server.persistence.enabled=false, or use a ReadWriteMany accessMode if you are keeping it for another reason." }}
{{- end }}
strategy:
type: Recreate
{{- end }}
selector:
matchLabels:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: server
template:
metadata:
labels:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: server
spec:
# The server stops gRPC before draining HTTP, so that every CommandStream
# handler returns and releases its agent's presence claim. A claim left
# behind outlives the pod for its 30s TTL, and during that window other
# replicas dispatch commands to a process that has exited — surfacing to
# the operator as "agent offline" on an agent that is perfectly healthy.
# 10s for gRPC plus 10s for the HTTP drain, with headroom.
terminationGracePeriodSeconds: {{ .Values.server.terminationGracePeriodSeconds }}
{{- if .Values.imagePullSecrets }}
imagePullSecrets:
{{- toYaml .Values.imagePullSecrets | nindent 8 }}
{{- end }}
{{- if or .Values.redis.enabled .Values.mongo.enabled }}
# Wait for the dependencies this chart deploys to be reachable,
# approximating compose's `depends_on: condition: service_healthy`. An
# external Redis or Mongo is assumed to be up already — waiting on one
# would only turn someone else's outage into a stuck pod.
initContainers:
{{- if .Values.redis.enabled }}
- name: wait-for-redis
image: busybox:1.36
command:
- sh
- -c
- |
until nc -z {{ .Release.Name }}-redis {{ .Values.redis.port }}; do
echo "waiting for redis..."; sleep 2;
done
{{- end }}
{{- if .Values.mongo.enabled }}
- name: wait-for-mongo
image: busybox:1.36
command:
- sh
- -c
- |
until nc -z {{ .Release.Name }}-mongo {{ .Values.mongo.port }}; do
echo "waiting for mongo..."; sleep 2;
done
{{- end }}
{{- end }}
containers:
- name: server
image: "{{ .Values.server.image.repository }}:{{ .Values.server.image.tag }}"
ports:
- containerPort: {{ .Values.server.service.httpPort }}
- containerPort: {{ .Values.server.service.grpcPort }}
env:
{{- include "vantage.server.env" . | nindent 12 }}
{{- if .Values.server.migrationJob.enabled }}
# Schema setup ran in the pre-upgrade Job. Pods that repeated it
# would race each other, and the rename migration is not a race
# that tolerates a loser.
- name: VANTAGE_SKIP_MIGRATIONS
value: "true"
{{- end }}
# Liveness never touches Mongo or Redis: restarting every pod cannot
# fix a database outage, and each restart drops every agent command
# stream and console session it was carrying. Readiness does check
# both, so a pod that cannot serve leaves the Service and stays up.
startupProbe:
httpGet:
path: /healthz
port: {{ .Values.server.service.httpPort }}
periodSeconds: 5
# Generous: without the migration Job this pod runs every migration
# before it listens, and the rename has a ten-minute budget.
failureThreshold: 150
livenessProbe:
httpGet:
path: /healthz
port: {{ .Values.server.service.httpPort }}
periodSeconds: 20
failureThreshold: 3
readinessProbe:
httpGet:
path: /readyz
port: {{ .Values.server.service.httpPort }}
periodSeconds: 10
failureThreshold: 3
{{- if .Values.server.persistence.enabled }}
# Nothing in the server writes here any more — workflow logs moved to
# MongoDB so that every replica can read and write them. The mount
# remains only so an operator upgrading from a file-log release can
# still reach the old files before turning persistence off.
volumeMounts:
- name: server-data
mountPath: /data
volumes:
- name: server-data
persistentVolumeClaim:
claimName: {{ .Release.Name }}-server-data
{{- end }}
---
{{- if gt (int .Values.server.replicaCount) 1 }}
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: {{ .Release.Name }}-server
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: server
spec:
# Agents reconnect on their own, but a drain that took every replica at once
# would disconnect every agent in the fleet simultaneously and stall every
# workflow run in flight.
minAvailable: 1
selector:
matchLabels:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: server
---
{{- end }}
apiVersion: v1
kind: Service
metadata:
name: {{ .Release.Name }}-server
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: server
spec:
type: {{ .Values.server.service.type }}
selector:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: server
ports:
- name: http
port: {{ .Values.server.service.httpPort }}
targetPort: {{ .Values.server.service.httpPort }}
- name: grpc
port: {{ .Values.server.service.grpcPort }}
targetPort: {{ .Values.server.service.grpcPort }}
+80
View File
@@ -0,0 +1,80 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: {{ .Release.Name }}-web
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: web
spec:
# web holds no per-process state: sessions live in Redis and every request is
# proxied to the server. It is the one component here that scales freely.
replicas: {{ .Values.web.replicaCount }}
selector:
matchLabels:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: web
template:
metadata:
labels:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: web
spec:
{{- if .Values.imagePullSecrets }}
imagePullSecrets:
{{- toYaml .Values.imagePullSecrets | nindent 8 }}
{{- end }}
initContainers:
- name: wait-for-server
image: busybox:1.36
command:
- sh
- -c
- |
until nc -z {{ .Release.Name }}-server {{ .Values.server.service.httpPort }}; do
echo "waiting for server..."; sleep 2;
done
containers:
- name: web
image: "{{ .Values.web.image.repository }}:{{ .Values.web.image.tag }}"
ports:
- containerPort: {{ .Values.web.service.port }}
env:
- name: API_URL
value: {{ tpl .Values.web.env.apiUrl . | quote }}
# /healthz is served by this Next process; /api is rewritten to the
# server, so a probe there would report the backend's health and keep
# passing while this pod was wedged.
startupProbe:
httpGet:
path: /healthz
port: {{ .Values.web.service.port }}
periodSeconds: 3
failureThreshold: 20
livenessProbe:
httpGet:
path: /healthz
port: {{ .Values.web.service.port }}
periodSeconds: 20
failureThreshold: 3
readinessProbe:
httpGet:
path: /healthz
port: {{ .Values.web.service.port }}
periodSeconds: 10
failureThreshold: 3
---
apiVersion: v1
kind: Service
metadata:
name: {{ .Release.Name }}-web
labels:
{{- include "vantage.labels" . | nindent 4 }}
app.kubernetes.io/component: web
spec:
type: {{ .Values.web.service.type }}
selector:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: web
ports:
- port: {{ .Values.web.service.port }}
targetPort: {{ .Values.web.service.port }}
+111
View File
@@ -0,0 +1,111 @@
# Default values for the vantage chart.
redis:
enabled: true
addr: ""
image:
repository: redis
tag: "8"
persistence:
enabled: true
size: 1Gi
storageClass: ""
accessMode: ReadWriteOnce
port: 6379
auth:
username: ""
password: ""
existingSecret: ""
usernameKey: username
passwordKey: password
mongo:
enabled: true
image:
repository: mongo
tag: "7"
persistence:
enabled: true
size: 5Gi
storageClass: ""
accessMode: ReadWriteOnce
port: 27017
guacd:
image:
repository: docker.io/guacamole/guacd
tag: "1.6.0"
service:
type: ClusterIP
port: 4822
server:
replicaCount: 1
# Must exceed the server's own stop sequence (10s gRPC GracefulStop + 10s
# HTTP drain) or the kubelet SIGKILLs mid-shutdown, which is exactly the
# abrupt exit that leaves agent presence claims stranded in Redis.
terminationGracePeriodSeconds: 30
migrationJob:
enabled: true
backoffLimit: 0
activeDeadlineSeconds: 900
image:
repository: gitea.hostxtra.co.uk/mrhid6/vantage/server
tag: latest
service:
type: ClusterIP
httpPort: 8080
grpcPort: 9090
env:
mongoUri: "mongodb://{{ .Release.Name }}-mongo:27017/vantage"
grpcHost: "{{ .Release.Name }}-server:9090"
keyEncryptionKey: ""
appRootLabel: vantage
proxyAdvertiseHost: "{{ .Release.Name }}-server"
proxyListenHost: "0.0.0.0"
persistence:
enabled: false
size: 1Gi
storageClass: ""
accessMode: ReadWriteOnce
hostPath: /data
web:
replicaCount: 1
image:
repository: gitea.hostxtra.co.uk/mrhid6/vantage/web
tag: latest
service:
type: ClusterIP
port: 3000
env:
apiUrl: "http://{{ .Release.Name }}-server:8080"
ingress:
enabled: false
className: traefik
entrypoint: websecure
annotations: {}
web:
host: ""
extraHosts: []
api:
enabled: false
paths:
- /api
- /auth
- /update
- /install
- /update.ps1
- /install.ps1
grpc:
enabled: true
host: ""
annotations: {}
tls:
enabled: true
secretName: ""
grpcSecretName: ""
certResolver: ""
imagePullSecrets: []
@@ -5,7 +5,7 @@
# host:port agents dial for gRPC. No default; boot fails without it.
# Must be reachable from managed servers. Use the public host, port 9090.
GRPC_HOST=192.168.1.250:9090
GRPC_HOST=vantage.yourdomain.com:9090
# 64-char hex (32 bytes) for AES-256-GCM. Required for private keys,
# secrets, OIDC secrets, RDP/VNC credentials.
@@ -18,4 +18,3 @@ KEY_ENCRYPTION_KEY=
MONGO_URI=mongodb://mongo:27017/vantage
# Where workflow run logs are written inside the server container.
# VANTAGE_WORKFLOW_LOG_DIR=/data/workflow-logs
@@ -45,9 +45,8 @@ services:
GRPC_PORT: "9090"
HTTP_PORT: "8080"
KEY_ENCRYPTION_KEY: ${KEY_ENCRYPTION_KEY:-}
VANTAGE_WORKFLOW_LOG_DIR: ${VANTAGE_WORKFLOW_LOG_DIR:-}
GUACD_ADDR: guacd:4822
APP_ROOT_LABEL: vantage
PROXY_ADVERTISE_HOST: server
depends_on:
redis:
condition: service_healthy
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,209 @@
# Agent-relayed console proxy
Date: 2026-07-29
Status: approved, not yet implemented
## Problem
`consoleTunnel` builds guacamole parameters from `srv.IPAddress` and hands them
to guacd, which then dials the target itself. On a self-hosted deployment the
control plane and the managed servers share a network, so that works. On Vantage
Cloud they do not: guacd runs on the cloud host and the customer's server is on
an RFC1918 address behind their NAT. Every cloud console session to a private
address fails, for SSH, RDP and VNC alike.
Agents already hold an outbound gRPC connection to the control plane. The fix is
to carry the console's TCP bytes over that existing path rather than asking guacd
to route somewhere it cannot reach.
## Decisions
**Self-relay only.** The agent relays to its own host and nowhere else. It is
never told a hostname; the host is hardcoded to `127.0.0.1` on the agent side and
only the port comes from the server. A jump-host mode (reaching agentless devices
through a neighbouring agent) was rejected: it would give an agent the power to
dial arbitrary addresses on the customer's LAN, and the console today can only
target servers that run an agent anyway.
**A dedicated bidirectional RPC, one stream per TCP connection.** Multiplexing
console bytes onto the existing `CommandStream` was rejected — that stream
already carries control commands and workflow stdout, and an RDP framebuffer
would introduce head-of-line blocking against key sync and step output. A
separate stream also gets connection lifetime, flow control and close semantics
for free instead of needing a hand-rolled connection-ID demux.
**Always proxy, both deployments.** Direct dial is deleted rather than kept as a
self-hosted fast path or a fallback. One code path means one tested code path,
and the cloud path is the one no developer can reproduce locally. A
try-direct-then-fall-back design was rejected outright: it puts a timeout in
front of every private-network session and makes "which path did this session
use" unanswerable from the audit log.
The cost is that the console now requires a live agent, where a self-hosted
deployment could previously reach a server whose agent was down. In practice an
offline agent almost always means an offline host, and the failure is now an
immediate, explicit refusal instead of a hang.
## Architecture
Three parties rendezvous on a single `proxy_id`. Neither guacd nor the agent
changes which direction it dials: guacd still makes an outbound TCP connection,
the agent still only connects outbound to the control plane.
```
consoleTunnel (server)
1. proxy.Open(instance, server_id, port) -> proxy_id + ephemeral listener :N
2. push OpenProxyCmd{proxy_id, port} down the existing CommandStream
3. agent dials 127.0.0.1:port locally, then opens ProxyStream and sends
ProxyOpen{server_id, agent_token, proxy_id}
4. guacd dials PROXY_ADVERTISE_HOST:N (the params it was handed in step 1)
5. registry holds both halves -> io.Copy in both directions
6. either side EOFs -> close listener, close stream, drop the registry entry
```
Steps 3 and 4 race, so a registry entry has two slots and starts piping when the
second one arrives. Both waits share a single 10 second deadline; expiry closes
everything and frees the entry.
The agent dials locally *before* opening the stream, so a refused connection
arrives as an explicit `ProxyClose{reason}` rather than as a hang.
`BuildGuacParams` stops reading `srv.IPAddress` and takes the relay host and port
instead. `IPAddress` remains in use for display and for monitors.
## Wire protocol
Additive only; no existing message changes shape.
```protobuf
rpc ProxyStream(stream ProxyClientMsg) returns (stream ProxyServerMsg);
message OpenProxyCmd { // ServerCommand oneof field 8
string proxy_id = 1;
uint32 port = 2;
}
message ProxyClientMsg {
oneof payload {
ProxyOpen open = 1; // first message only
bytes data = 2;
ProxyClose close = 3;
}
}
message ProxyOpen { string server_id = 1; string agent_token = 2; string proxy_id = 3; }
message ProxyServerMsg { oneof payload { bytes data = 1; ProxyClose close = 2; } }
message ProxyClose { string reason = 1; }
```
Two implementation facts about this repo shape the above. The `pb` packages are
**hand-written Go, not protoc output** — `vantage.proto` is documentation, and
both `server/internal/grpc/pb` and `agent/internal/grpc/pb` are edited by hand
and kept in sync manually. And the registered codec is JSON, so a `bytes` field
travels as a base64 string: roughly 33% overhead on relayed traffic. That is
accepted rather than fixed here, because introducing a second codec for one RPC
is a larger change than this feature warrants. Relay chunks are 32 KiB.
## Security
**The agent only ever dials `127.0.0.1`.** The port is the only field it takes
from the server; the host is hardcoded agent-side. A compromised control plane
cannot use an agent to reach anything else on the customer's network. This is the
strongest property in the design and the reason self-relay was chosen.
**`proxy_id` is 32 random bytes, single-use and scoped.** On `ProxyOpen` the
server checks three things together: the agent token hash matches that
`server_id`, the `proxy_id` exists in the registry, and the entry's `server_id`
and `instance_id` match the authenticated agent. Any mismatch closes the stream
without revealing which check failed.
**The listener is the exposed surface and is narrowed four ways.** It binds an
ephemeral port; it lives at most 10 seconds unclaimed; it accepts exactly one
connection and closes immediately afterwards; and the accepted connection's
remote address must resolve to a host named in `GUACD_ADDR`. Without that last
check, any other container on the Docker network could claim the session during
the window.
**Agent-offline is refused early.** `consoleConnect` checks
`srv.Status == "active"` and returns 409 `agent_offline`, rather than letting the
browser open a WebSocket that dies on a deadline.
**Audit.** `console.opened` gains the relay port and `proxy_id`. A relay that
expires or is refused writes `console.proxy_failed` with a reason, so a failed
console session stops being invisible.
Credentials are unchanged. Private keys and RDP passwords travel from the server
to guacd inside the guacamole handshake and never reach the agent. The SSH and
RDP sessions are negotiated end-to-end between guacd and the target daemon, so
the agent relays bytes it cannot read.
## Components
New, server:
| Unit | Responsibility |
| --- | --- |
| `server/internal/proxy/registry.go` | `Open`, `AttachAgent`, `AttachTCP`, expiry sweep. Pure state — no net, no gRPC, testable alone |
| `server/internal/proxy/session.go` | One relay: listener, deadline, the `io.Copy` pair, teardown-once |
| `server/internal/grpc/proxystream.go` | The `ProxyStream` handler: authenticate, then hand the stream to the registry. No relay logic of its own |
New, agent:
| Unit | Responsibility |
| --- | --- |
| `agent/internal/proxy/proxy.go` | `Open(ctx, client, proxyID, port)` — dial loopback, open the stream, pump bytes. No build tags; Linux and Windows share it |
Changed:
- `proto/vantage/v1/vantage.proto`, and both generated pb trees
- `server/internal/services/console.go``BuildGuacParams(srv, relayHost, relayPort, …)`
- `server/internal/api/console.go` — offline pre-check in `consoleConnect`; open the relay before the guacd handshake in `consoleTunnel` and close it in `OnDisconnect`
- `agent/internal/sync/sync.go` — handle `OpenProxyCmd`, one goroutine per proxy
- `deploy/docker-compose.yml`, `deploy/docker-compose.site.yml``PROXY_ADVERTISE_HOST=server`
Two new optional environment variables on the server: `PROXY_ADVERTISE_HOST`
(default `server`, the name guacd resolves the control plane by) and
`PROXY_LISTEN_HOST` (default `0.0.0.0`).
Nothing new is opened on the customer's firewall — the relay rides the agent's
existing outbound gRPC connection.
`docsite/docs/reference/ports-and-networking.md` and
`docsite/docs/vantage/browser-console.md` must say so, and must state the new
requirement that the agent be online.
A secondary benefit beyond cloud: a VNC or RDP service bound only to `127.0.0.1`
is now reachable, where a direct dial from guacd never could be.
## Failure modes
| Failure | Behaviour |
| --- | --- |
| Agent offline at connect | 409 `agent_offline` from `consoleConnect`, before any WebSocket is opened |
| Agent never opens the stream | 10s deadline; listener closed; `console.proxy_failed{reason:"agent_timeout"}`; WebSocket closed with a message the UI surfaces |
| Local dial refused (daemon down, wrong port) | `ProxyClose{reason}` relayed up as the same audit event, reason `dial_refused` |
| guacd never dials | Same deadline path, reason `guacd_timeout` |
| Bad token, unknown or foreign `proxy_id` | Stream closed with no detail leaked; `console.proxy_failed{reason:"rejected"}` |
| Agent process dies mid-session | Stream EOF, relay torn down, console shows a disconnect |
| CommandStream reconnects mid-session | No effect on live sessions — the relay is on its own stream. Only a new `OpenProxyCmd` needs the control stream |
Teardown is guarded by `sync.Once` on both sides: both `io.Copy` goroutines
finish, and whichever finishes second must not double-close.
## Testing
Written test-first.
- `server/internal/proxy/registry_test.go` — the two halves pair in either
order; expiry frees the entry; a second claim on a used `proxy_id` is
rejected; a mismatched `instance_id` is rejected. No network.
- `server/internal/proxy/session_test.go` — two `net.Pipe` halves; bytes flow
both ways; EOF in each direction tears down; double-close is safe.
- `server/internal/grpc/proxystream_test.go` — the authentication matrix: valid,
wrong token, unknown `proxy_id`, `proxy_id` belonging to another instance.
- `agent/internal/proxy` — a refused dial emits `ProxyClose`; the happy path
echoes bytes.
- End-to-end in `server`: a fake agent plus a `net.Listen` echo server, asserting
bytes traverse listener → registry → stream → echo and back. This is the test
that would have caught the original bug.
Manual verification, in this order: self-hosted SSH (proves no regression),
cloud SSH to a private-network host, cloud RDP to a Windows agent.
@@ -50,7 +50,6 @@ GRPC_HOST=vantage.example.com:9090
KEY_ENCRYPTION_KEY=
# Optional: where workflow run logs are written inside the container.
VANTAGE_WORKFLOW_LOG_DIR=/data/workflow-logs
```
Generate the encryption key:
@@ -13,12 +13,16 @@ it is absent.
| -------------------------- | --------------- | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `GRPC_HOST` | **yes** | | The `host:port` agents dial. Boot fails without it. There is deliberately no fallback to the web host: that would hand every agent a port that does not speak gRPC |
| `MONGO_URI` | no | `mongodb://localhost:27017` | The database name is taken from the URI path, falling back to `vantage`. There is no separate `MONGO_DB` |
| `REDIS_ADDR` | no | `localhost:6379` | Sessions only |
| `REDIS_ADDR` | no | `localhost:6379` | Sessions, and the bus that routes agent commands between server replicas. Every replica must point at the **same** Redis |
| `REDIS_USERNAME` | no | | Redis 6+ ACL user. Leave empty against a legacy `requirepass` instance, which authenticates with the password alone |
| `REDIS_PASSWORD` | no | | Leave empty for an unauthenticated Redis. Both of these exist so an install can use a managed Redis rather than the bundled one |
| `KEY_ENCRYPTION_KEY` | yes in practice | | 64 hex characters (32 bytes) for AES-256-GCM. Required for private keys, vault secrets, OIDC client secrets and console credentials |
| `GITEA_HOST` | yes | `gitea.example.com` | Used to build the install scripts and agent download URLs. The default is a placeholder that will not resolve |
| `GUACD_ADDR` | no | `guacd:4822` | The [browser console](../vantage/browser-console.md) daemon |
| `PROXY_ADVERTISE_HOST` | no | `server` | The hostname **guacd** uses to reach the control plane's console relay. Wrong here and every console session fails at connect with guacd unable to resolve the relay |
| `PROXY_LISTEN_HOST` | no | `0.0.0.0` | Interface the ephemeral relay listeners bind. Narrow it only if guacd shares a known interface |
| `POD_IP` | no | | Kubernetes only, set by the Helm chart from the downward API. Overrides `PROXY_ADVERTISE_HOST`, because a console relay belongs to one replica and a Service address names all of them |
| `APP_ROOT_LABEL` | no | `vantage` | The app root label for the host and session organisation guard |
| `VANTAGE_WORKFLOW_LOG_DIR` | no | | Where workflow run logs are written |
:::danger `KEY_ENCRYPTION_KEY` has no recovery path
It encrypts SSH private keys, vault secrets, OIDC client secrets and console
+12 -5
View File
@@ -24,7 +24,8 @@ flowchart LR
W --> S["server :8080"]
A["Agent on a managed server"] -->|"gRPC/TLS :9090, outbound"| S
S --> G["guacd :4822"]
G -->|"SSH / RDP / VNC"| T["Target machine"]
G -->|"relayed over the :9090 stream"| A
A -->|"SSH / RDP / VNC, loopback"| T["Target machine (same host as agent)"]
```
Two things are worth reading off that diagram.
@@ -33,9 +34,13 @@ Two things are worth reading off that diagram.
NAT is not an obstacle. The only requirement is that the machine can reach
`GRPC_HOST`.
**The console does not use the agent.** guacd connects directly to the target on
the protocol port. A machine reachable only by its agent behind NAT, on a
private subnet cannot be consoled, even though every other feature works.
**The console rides the agent's connection too.** guacd never dials the target
directly; the server pushes a command down the agent's existing outbound gRPC
stream on `9090`, and the agent relays the protocol traffic from its own
loopback. No route from the control plane to the target's address is needed,
and no new inbound port opens on the target — the same connection that carries
key sync carries console traffic. This is what makes the console work for a
machine behind NAT on a private subnet, as long as its agent is online.
## What to open
@@ -49,7 +54,9 @@ private subnet cannot be consoled, even though every other feature works.
- `gitea.hostxtra.co.uk`, for agent releases and version checks.
- Anything a server-run [monitor](../vantage/monitors.md) checks.
- SMTP, if you use an SMTP notification channel.
- Protocol ports on machines you intend to console.
No route to the machines you intend to console is needed — that traffic rides
the agent's existing outbound `9090` connection instead.
### Outbound from a managed machine
+1 -1
View File
@@ -86,7 +86,7 @@ instantaneous.
| Connects, then closes at once | guacd unreachable. Check `GUACD_ADDR` and that the container is running |
| SSH rejects the key | The stored key has no private half, or is not on the target |
| RDP fails on retry | Credentials are single-use and consumed at tunnel open enter them again |
| Hangs at "connecting" | The **control plane** cannot reach the target on the protocol port. The agent's reachability is irrelevant here |
| Hangs, then disconnects | The agent never claimed the relay, nothing is listening on the protocol port on the target's own loopback address, or guacd never dialled in time. Check the audit log for `console.proxy_failed` — its reason (`agent_timeout`, `dial_refused`, `guacd_timeout`, `rejected`) names which |
| Fails only in production | The reverse proxy is not forwarding WebSocket upgrade headers |
## Monitors report down when the service is up
+9 -6
View File
@@ -15,17 +15,20 @@ to a **guacd** daemon and manages credentials around it.
- `guacd` running and reachable from the server. The bundled Compose stack
includes it; `GUACD_ADDR` defaults to `guacd:4822`.
- `KEY_ENCRYPTION_KEY` set, since every credential involved is stored encrypted.
- Network reachability **from the control plane to the target** on the protocol
port. This is the one part of Vantage that is not agent-mediated: guacd
connects directly, so a machine reachable only by its agent cannot be
consoled.
- The target's **agent must be online**. Console traffic is relayed over the
agent's existing outbound connection, so the control plane never needs a route
to the server's address — but it does need the agent.
- No inbound port on the target, beyond what the protocol already listens on
locally. A service bound only to `127.0.0.1` works, because the agent dials
loopback on the target itself.
## Opening a session
From a server's page, choose **Console**. Then:
1. The UI calls `POST /api/console/connect`, which mints a **one-time** session
token.
token. If the target's agent is not connected, this fails immediately with
`409 agent_offline` rather than hanging.
2. The browser opens a WebSocket to `GET /api/console/tunnel` with that token.
3. The server marks the token consumed atomically, so a second use cannot
race and proxies the connection to guacd.
@@ -67,4 +70,4 @@ keystroke log. If you need that, it has to come from the target machine.
| Connects then closes immediately | guacd unreachable check `GUACD_ADDR` and that the container is up |
| SSH refuses the key | The stored key has no private half, or is not in the target's `authorized_keys` |
| RDP fails on a fresh credential | Credentials are consumed on open; a retry needs them entered again |
| Hangs at connecting | The control plane cannot reach the target on the protocol port |
| Hangs, then disconnects | The agent never claimed the relay, nothing is listening on the protocol port on the target's own loopback address, or guacd never dialled in time. Check the audit log for `console.proxy_failed` — its reason (`agent_timeout`, `dial_refused`, `guacd_timeout`, `rejected`) names which |
+32
View File
@@ -14,6 +14,7 @@ service Vantage {
rpc ReportChecks(ReportChecksRequest) returns (ReportChecksResponse);
// Bidirectional stream: agent sends auth once, server pushes commands.
rpc CommandStream(stream AgentMessage) returns (stream ServerCommand);
rpc ProxyStream(stream ProxyClientMsg) returns (stream ProxyServerMsg);
}
message RegisterRequest {
@@ -180,6 +181,7 @@ message ServerCommand {
ApplyUpdatesCmd apply_updates = 5;
RunStepCmd run_step = 6;
CleanupWorkspaceCmd cleanup_workspace = 7;
OpenProxyCmd open_proxy = 8;
}
}
@@ -228,3 +230,33 @@ message StepOutputChunk {
bytes data = 3;
bool eof = 4;
}
// OpenProxyCmd tells the agent to dial 127.0.0.1:port locally and relay that
// connection back over a fresh ProxyStream identified by proxy_id.
message OpenProxyCmd {
string proxy_id = 1;
uint32 port = 2;
}
message ProxyOpen {
string server_id = 1;
string agent_token = 2;
string proxy_id = 3;
}
message ProxyClose { string reason = 1; }
message ProxyClientMsg {
oneof payload {
ProxyOpen open = 1; // first message only
bytes data = 2;
ProxyClose close = 3;
}
}
message ProxyServerMsg {
oneof payload {
bytes data = 1;
ProxyClose close = 2;
}
}
+125 -22
View File
@@ -2,12 +2,18 @@ package main
import (
"context"
"errors"
"log"
"net/http"
"os"
"os/signal"
"strings"
"syscall"
"time"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/api"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/auth"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/bus"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/db"
grpcserver "gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/grpc"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/monitorsched"
@@ -18,7 +24,21 @@ import (
func main() {
mongoURI := getEnv("MONGO_URI", "mongodb://localhost:27017")
if os.Getenv("GRPC_HOST") == "" {
// Two flags, so the schema work can be lifted out of the serving pods.
//
// Under Docker Compose neither is set and nothing changes: one process
// migrates and then serves. Under Kubernetes with more than one replica
// that is unsafe — every pod would run MigrateOrgToInstance at once, and
// renaming collections while a sibling reads them is not a race anyone
// wins. The chart therefore runs a pre-upgrade Job with MIGRATE_ONLY and
// starts the Deployment with SKIP_MIGRATIONS.
migrateOnly := boolEnv("VANTAGE_MIGRATE_ONLY")
skipMigrations := boolEnv("VANTAGE_SKIP_MIGRATIONS")
// Not required in migrate-only mode: that process never serves gRPC, and
// demanding it would put an agent-facing address in a Job that has no
// business knowing one.
if !migrateOnly && os.Getenv("GRPC_HOST") == "" {
log.Fatal("GRPC_HOST is required (host:port agents dial for gRPC)")
}
@@ -28,6 +48,24 @@ func main() {
}
log.Println("connected to MongoDB")
if migrateOnly || !skipMigrations {
runSchemaSetup()
} else {
log.Println("VANTAGE_SKIP_MIGRATIONS set: assuming migrations ran elsewhere")
}
if migrateOnly {
log.Println("VANTAGE_MIGRATE_ONLY set: schema setup complete, exiting")
return
}
serve()
}
// runSchemaSetup performs every write that must happen exactly once before the
// application serves: migrations, index builders and default step seeding. It
// is fatal on anything that would leave the schema half-moved.
func runSchemaSetup() {
// Migrations 0001 to 0003 still speak the pre-rename shape (orgs, org_id),
// so they must run before 0004 renames everything underneath them.
if err := services.RunMigrations(); err != nil {
@@ -87,34 +125,63 @@ func main() {
}
}
services.StartLogSweeper()
services.StartAuditSweeper()
}
func serve() {
redisAddr := getEnv("REDIS_ADDR", "localhost:6379")
if err := auth.InitRedis(redisAddr); err != nil {
redisUser := os.Getenv("REDIS_USERNAME")
redisPass := os.Getenv("REDIS_PASSWORD")
if err := auth.InitRedis(redisAddr, redisUser, redisPass); err != nil {
log.Fatalf("failed to connect to Redis: %v", err)
}
log.Println("connected to Redis")
go func() {
// The bus carries agent commands and step results between replicas. It is
// not optional even on a single-replica deployment: dispatch takes the same
// path either way, so the code exercised in production is the code
// exercised everywhere.
if err := bus.Init(redisAddr, redisUser, redisPass); err != nil {
log.Fatalf("failed to connect the message bus: %v", err)
}
log.Printf("message bus ready as node %s", bus.NodeID())
// Cancelled on SIGTERM/SIGINT. Everything below that takes a context — the
// housekeeping jobs, the leader lock — stops when the pod is asked to.
ctx, shutdown := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer shutdown()
stopGRPC, err := grpcserver.StartGRPC(9090)
if err != nil {
log.Fatalf("gRPC server error: %v", err)
}
// Everything below runs on exactly one replica at a time.
//
// These are cluster-singleton jobs, not per-pod work: N replicas would mean
// every monitor check firing N times, every incident notification delivered
// to the customer N times, every retention sweep deleting concurrently, and
// N reapers racing to purge the same instance. They share one lock rather
// than holding four, because they are one role — housekeeping — and
// splitting them would only spread that role across pods for no benefit.
bus.RunAsLeader(ctx, "housekeeping", func(jobCtx context.Context) {
services.StartLogSweeper(jobCtx)
services.StartAuditSweeper(jobCtx)
services.StartReaper(jobCtx)
monitorsched.Start(jobCtx)
ticker := time.NewTicker(2 * time.Minute)
defer ticker.Stop()
for range ticker.C {
if err := services.MarkOfflineServers(); err != nil {
log.Printf("mark offline error: %v", err)
for {
select {
case <-jobCtx.Done():
return
case <-ticker.C:
if err := services.MarkOfflineServers(); err != nil {
log.Printf("mark offline error: %v", err)
}
}
}
}()
go func() {
if err := grpcserver.StartGRPC(9090); err != nil {
log.Fatalf("gRPC server error: %v", err)
}
}()
monitorsched.Start(context.Background())
services.StartReaper(context.Background())
})
r := gin.New()
r.Use(gin.Recovery())
@@ -122,12 +189,37 @@ func main() {
r.Use(corsMiddleware())
api.RegisterRoutes(r)
log.Println("REST server listening on :8080")
if err := r.Run(":8080"); err != nil {
log.Fatalf("REST server error: %v", err)
srv := &http.Server{Addr: ":8080", Handler: r}
go func() {
log.Println("REST server listening on :8080")
if err := srv.ListenAndServe(); err != nil && !errors.Is(err, http.ErrServerClosed) {
log.Fatalf("REST server error: %v", err)
}
}()
<-ctx.Done()
log.Println("shutdown signal received")
// gRPC first, and this ordering is the point of the whole exercise. Stopping
// it runs each CommandStream handler's deferred release, which clears that
// agent's presence claim; until that happens another replica will keep
// dispatching commands to this process. Draining HTTP first would leave the
// claims held for the length of the drain.
stopGRPC()
drainCtx, cancelDrain := context.WithTimeout(context.Background(), httpDrainTimeout)
defer cancelDrain()
if err := srv.Shutdown(drainCtx); err != nil {
log.Printf("REST server shutdown: %v", err)
}
log.Println("shutdown complete")
}
// How long in-flight REST requests are given to finish. Console tunnels are
// long-lived WebSockets that will not end on their own, so this is a ceiling
// rather than a target; the relays behind them are already gone by this point.
const httpDrainTimeout = 10 * time.Second
func corsMiddleware() gin.HandlerFunc {
return func(c *gin.Context) {
c.Header("Access-Control-Allow-Origin", "*")
@@ -147,3 +239,14 @@ func getEnv(key, fallback string) string {
}
return fallback
}
// boolEnv reads a flag env var. Anything other than a recognised truthy value
// is false, so a typo leaves the safe default (migrate here, serve here) rather
// than silently skipping schema setup.
func boolEnv(key string) bool {
switch strings.ToLower(strings.TrimSpace(os.Getenv(key))) {
case "1", "true", "yes", "on":
return true
}
return false
}
+115 -5
View File
@@ -1,6 +1,9 @@
package api
import (
"errors"
"fmt"
"log"
"net"
"net/http"
"os"
@@ -33,6 +36,15 @@ func consoleConnect(c *gin.Context) {
return
}
if srv.Status != "active" {
c.JSON(http.StatusConflict, gin.H{
"error": "agent_offline",
"message": "The agent on this server is not connected. " +
"Console sessions are relayed by the agent, so it must be online.",
})
return
}
sess, err := services.CreateConsoleSession(auth.InstanceID(c), body.ServerID, body.Protocol, body.KeyID, actorFromCtx(c), c.ClientIP())
if err != nil {
c.JSON(http.StatusInternalServerError, gin.H{"error": err.Error()})
@@ -59,7 +71,7 @@ func consoleConnect(c *gin.Context) {
}
services.LogEvent(auth.InstanceID(c), "console.opened", actorFromCtx(c), srv.ServerID, "",
"console session opened ("+body.Protocol+")")
"console session opened ("+body.Protocol+", agent-relayed)")
c.JSON(http.StatusOK, gin.H{
"session_id": sess.SessionID,
@@ -76,56 +88,137 @@ func queryIntDefault(r *http.Request, key string, def int) int {
return v
}
// consoleTunnel upgrades the browser's WebSocket and joins it to guacd.
//
// Every branch here logs. That is deliberate and worth keeping: this handler
// spans four hops (session store, agent dispatch, relay announcement, guacd),
// any of which can fail, and the client is told the same near-useless thing by
// most of them — a 500 that guacamole then reports as an *upstream* error,
// naming the wrong hop entirely. Without a line per branch the only evidence a
// failure leaves is a GIN status code, and with several replicas you cannot
// even tell which process produced it.
//
// Lines are prefixed with the session ID so one attempt can be followed across
// pods, and the pod's own hostname so it is obvious which one served it.
func consoleTunnel(c *gin.Context) {
host, _ := os.Hostname()
token := c.Query("token")
sessionID, err := services.VerifySessionToken(token)
if err != nil {
log.Printf("console[%s]: reject: invalid session token: %v", host, err)
c.JSON(http.StatusUnauthorized, gin.H{"error": "invalid token"})
return
}
// Bound to the session from here on, so every later line correlates.
tlog := func(format string, args ...any) {
log.Printf("console[%s %s]: "+format, append([]any{host, sessionID}, args...)...)
}
tlog("tunnel opened by %s", actorFromCtx(c))
instanceID := auth.InstanceID(c)
sess, err := services.GetConsoleSession(instanceID, sessionID)
if err != nil {
tlog("reject: console session not found: %v", err)
c.JSON(http.StatusNotFound, gin.H{"error": "session not found"})
return
}
if actor := actorFromCtx(c); actor != sess.User {
tlog("reject: session belongs to %s, not %s", sess.User, actor)
c.JSON(http.StatusForbidden, gin.H{"error": "session belongs to another user"})
return
}
if err := services.ConsumeSessionToken(instanceID, sessionID); err != nil {
tlog("reject: token already consumed: %v", err)
c.JSON(http.StatusUnauthorized, gin.H{"error": "token already used"})
return
}
srv, err := services.GetServer(auth.InstanceID(c), sess.ServerID)
if err != nil {
tlog("reject: server %s not found: %v", sess.ServerID, err)
c.JSON(http.StatusNotFound, gin.H{"error": "server not found"})
return
}
tlog("server %s (%s), protocol %s", srv.ServerID, srv.Hostname, sess.Protocol)
var privKey, passphrase string
if sess.Protocol == "ssh" && sess.KeyID != "" {
privKey, err = services.GetPrivateKey(auth.InstanceID(c), sess.KeyID)
if err != nil {
tlog("reject: key %s has no private material: %v", sess.KeyID, err)
c.JSON(http.StatusBadRequest, gin.H{"error": "selected key has no private material"})
return
}
passphrase, _ = services.GetPassphrase(sess.KeyID)
tlog("ssh key %s loaded (passphrase=%t)", sess.KeyID, passphrase != "")
}
var rdpUser, rdpPass string
if sess.Protocol == "rdp" || sess.Protocol == "vnc" {
rdpUser, rdpPass, err = services.ConsumeConsoleRDPCreds(instanceID, sessionID)
if err != nil {
tlog("reject: could not consume %s credentials: %v", sess.Protocol, err)
c.JSON(http.StatusInternalServerError, gin.H{"error": "could not load credentials"})
return
}
tlog("%s credentials consumed (user=%t)", sess.Protocol, rdpUser != "")
}
gp, err := services.BuildGuacParams(srv, sess.Protocol, sess.SSHUsername, privKey, passphrase, rdpUser, rdpPass)
targetPort, err := services.TargetPort(srv, sess.Protocol)
if err != nil {
tlog("reject: no target port for %s: %v", sess.Protocol, err)
c.JSON(http.StatusBadRequest, gin.H{"error": err.Error()})
return
}
tlog("opening relay to %s:%d", srv.ServerID, targetPort)
relay, err := services.OpenConsoleProxy(instanceID, srv.ServerID, targetPort)
if err != nil {
if errors.Is(err, services.ErrAgentOffline) {
tlog("reject: agent offline")
c.JSON(http.StatusConflict, gin.H{"error": "agent_offline"})
return
}
// The client is deliberately told nothing specific, so this is the only
// place the real reason exists — a failed dispatch and a relay that was
// never announced are the same generic 500 to the browser.
tlog("reject: open relay: %v", err)
c.JSON(http.StatusInternalServerError, gin.H{"error": "could not open relay"})
return
}
tlog("relay %s ready at %s:%d", relay.ProxyID, relay.Host, relay.Port)
// guac.WebsocketServer.ServeHTTP returns before installing its
// OnDisconnect handler when the connect callback errors, which is exactly
// the path every relay failure this proxy introduces takes (the agent
// never claims it, dial_refused, guacd never dials, rejected). Emitting
// console.proxy_failed and ending the session here, unconditionally on
// teardown, is what makes those failures reach the audit log at all;
// OnDisconnect below only sees the rarer case of a session that was fully
// established and then failed.
defer func() {
relay.Close()
if reason := relay.Reason(); reason != "" {
tlog("relay %s ended: %s", relay.ProxyID, reason)
services.LogEvent(instanceID, "console.proxy_failed", actorFromCtx(c), srv.ServerID, "",
fmt.Sprintf("console relay failed: %s (proxy_id=%s, port=%d)", reason, relay.ProxyID, relay.Port))
} else {
tlog("relay %s closed cleanly", relay.ProxyID)
}
_ = services.EndConsoleSession(instanceID, sessionID)
tlog("tunnel finished")
}()
services.LogEvent(instanceID, "console.proxy_opened", actorFromCtx(c), srv.ServerID, "",
fmt.Sprintf("console relay opened (proxy_id=%s, port=%d)", relay.ProxyID, relay.Port))
gp, err := services.BuildGuacParams(sess.Protocol, sess.SSHUsername, privKey, passphrase,
rdpUser, rdpPass, relay.Host, relay.Port)
if err != nil {
tlog("reject: build guacd parameters: %v", err)
c.JSON(http.StatusBadRequest, gin.H{"error": err.Error()})
return
}
@@ -145,24 +238,41 @@ func consoleTunnel(c *gin.Context) {
config.OptimalScreenHeight = queryIntDefault(r, "height", 768)
config.OptimalResolution = queryIntDefault(r, "dpi", 96)
// Resolution is logged separately from the dial: a headless guacd
// Service returns pod addresses, and which one was picked is the
// difference between "guacd refused" and "we called the wrong guacd".
addr, err := net.ResolveTCPAddr("tcp", guacdAddr)
if err != nil {
tlog("guacd: resolve %s: %v", guacdAddr, err)
return nil, err
}
tlog("guacd: dialling %s (%s)", guacdAddr, addr.String())
conn, err := net.DialTCP("tcp", nil, addr)
if err != nil {
tlog("guacd: dial %s: %v", addr.String(), err)
return nil, err
}
// The handshake is where guacd connects onward to the relay, so a
// failure here is guacd reporting it could not reach %s:%d — the hop
// that has been hardest to see from either end.
stream := guac.NewStream(conn, guac.SocketTimeout)
if err := stream.Handshake(config); err != nil {
tlog("guacd: handshake for %s to relay %s:%d: %v",
gp.Protocol, relay.Host, relay.Port, err)
return nil, err
}
tlog("guacd: tunnel established (%s %dx%d)",
gp.Protocol, config.OptimalScreenWidth, config.OptimalScreenHeight)
return guac.NewSimpleTunnel(stream), nil
}
// Teardown (proxy_failed audit + EndConsoleSession) lives in the deferred
// func above, not here: this only fires once a tunnel was actually
// established, and letting both paths log would double the audit event.
wsServer := guac.NewWebsocketServer(connect)
wsServer.OnDisconnect = func(id string, r *http.Request, t guac.Tunnel) {
_ = services.EndConsoleSession(instanceID, sessionID)
}
tlog("serving websocket")
wsServer.ServeHTTP(c.Writer, c.Request)
tlog("websocket returned")
}
+5 -3
View File
@@ -20,6 +20,9 @@ func actorFromCtx(c *gin.Context) string {
}
func RegisterRoutes(r *gin.Engine) {
r.GET("/healthz", handleHealthz)
r.GET("/readyz", handleReadyz)
r.GET("/install", handleInstallScript)
r.GET("/install.ps1", handleInstallScriptWindows)
r.GET("/update", handleUpdateScript)
@@ -139,7 +142,6 @@ func newServer(c *gin.Context) {
}
services.LogEvent(auth.InstanceID(c), "server.created", actorFromCtx(c), s.ServerID, "", "pre-registration token issued")
host := publicHostFromRequest(c)
installCmd := fmt.Sprintf(
@@ -412,7 +414,7 @@ if [ -z "$LATEST" ]; then
fi
VERSION="${LATEST#agent/}"
LATEST_ENCODED="${LATEST/\
LATEST_ENCODED="${LATEST/\//%%2F}"
BINARY_URL="https://${GITEA_HOST}/mrhid6/vantage/releases/download/${LATEST_ENCODED}/vantage-agent-linux-${ARCH}"
CHECKSUM_URL="https://${GITEA_HOST}/mrhid6/vantage/releases/download/${LATEST_ENCODED}/checksums.txt"
@@ -519,7 +521,7 @@ if [ -z "$LATEST" ]; then
fi
VERSION="${LATEST#agent/}"
LATEST_ENCODED="${LATEST/\
LATEST_ENCODED="${LATEST/\//%%2F}"
BINARY_URL="https://${GITEA_HOST}/mrhid6/vantage/releases/download/${LATEST_ENCODED}/vantage-agent-linux-${ARCH}"
CHECKSUM_URL="https://${GITEA_HOST}/mrhid6/vantage/releases/download/${LATEST_ENCODED}/checksums.txt"
+61
View File
@@ -0,0 +1,61 @@
package api
import (
"context"
"net/http"
"time"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/auth"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/db"
"github.com/gin-gonic/gin"
)
// The two probes answer different questions on purpose.
//
// /healthz is liveness: the process is up and serving. It touches nothing
// external, because a Mongo outage must not make Kubernetes restart every
// server pod — a restart loop cannot fix someone else's database, and it
// destroys every open command stream and console session on the way.
//
// /readyz is readiness: this pod can serve a request end to end, which needs
// both Mongo and Redis. A failing readiness probe pulls the pod out of the
// Service and leaves it running, which is the behaviour that matters during a
// dependency blip.
//
// Both sit outside /api, so neither the session middleware nor the licence
// gate applies. Neither reveals anything beyond up or down.
const probeTimeout = 2 * time.Second
func handleHealthz(c *gin.Context) {
c.JSON(http.StatusOK, gin.H{"status": "ok"})
}
func handleReadyz(c *gin.Context) {
ctx, cancel := context.WithTimeout(c.Request.Context(), probeTimeout)
defer cancel()
checks := gin.H{"mongo": "ok", "redis": "ok"}
ready := true
if db.Client == nil {
checks["mongo"] = "not initialised"
ready = false
} else if err := db.Client.Ping(ctx, nil); err != nil {
checks["mongo"] = "unreachable"
ready = false
}
if err := auth.PingRedis(ctx); err != nil {
checks["redis"] = "unreachable"
ready = false
}
status := http.StatusOK
state := "ok"
if !ready {
status = http.StatusServiceUnavailable
state = "unready"
}
c.JSON(status, gin.H{"status": state, "checks": checks})
}
+41 -25
View File
@@ -5,7 +5,6 @@ import (
"fmt"
"io"
"net/http"
"os"
"regexp"
"strconv"
"strings"
@@ -50,23 +49,42 @@ func getServerRunLog(c *gin.Context) {
c.JSON(http.StatusBadRequest, gin.H{"error": "invalid id"})
return
}
path := services.ServerRunLogPath(runID, serverID)
b, err := os.ReadFile(path)
if err != nil {
if !services.HasServerRunLog(runID, serverID) {
c.JSON(http.StatusNotFound, gin.H{"error": "no logs"})
return
}
c.Data(http.StatusOK, "text/plain; charset=utf-8", b)
c.Header("Content-Type", "text/plain; charset=utf-8")
c.Status(http.StatusOK)
// Streamed in pages rather than read whole. A log capped at 200k lines is
// tens of megabytes, and holding that in memory per concurrent download is
// how one curious user takes a pod down.
var after int64
for {
lines, last, err := services.ReadServerRunLog(runID, serverID, after, logPageSize)
if err != nil || len(lines) == 0 {
return
}
for _, l := range lines {
_, _ = c.Writer.WriteString(l)
_, _ = c.Writer.WriteString("\n")
}
c.Writer.Flush()
after = last
}
}
// logPageSize bounds one read of the log store. Large enough that a normal log
// is one or two queries, small enough that no single response buffers much.
const logPageSize = 2000
func streamServerRunLog(c *gin.Context) {
runID, serverID := c.Param("runId"), c.Param("serverId")
if !uuidLike.MatchString(runID) || !uuidLike.MatchString(serverID) {
c.JSON(http.StatusBadRequest, gin.H{"error": "invalid id"})
return
}
path := services.ServerRunLogPath(runID, serverID)
c.Writer.Header().Set("Content-Type", "text/event-stream")
c.Writer.Header().Set("Cache-Control", "no-cache")
c.Writer.Header().Set("Connection", "keep-alive")
@@ -78,29 +96,27 @@ func streamServerRunLog(c *gin.Context) {
return
}
var offset int64
// The cursor is a sequence number now, not a byte offset: lines come from
// the log store, which any pod can read, rather than from a file only this
// one has.
var after int64
sendNew := func() {
f, err := os.Open(path)
if err != nil {
return
}
defer f.Close()
if _, err := f.Seek(offset, 0); err != nil {
return
}
buf := make([]byte, 32*1024)
for {
n, _ := f.Read(buf)
if n <= 0 {
break
lines, last, err := services.ReadServerRunLog(runID, serverID, after, logPageSize)
if err != nil || len(lines) == 0 {
return
}
offset += int64(n)
for _, line := range splitSSE(buf[:n]) {
_, _ = c.Writer.WriteString("data: " + line + "\n")
after = last
for _, line := range lines {
for _, part := range splitSSE([]byte(line)) {
_, _ = c.Writer.WriteString("data: " + part + "\n")
}
_, _ = c.Writer.WriteString("\n")
}
_, _ = c.Writer.WriteString("\n")
flusher.Flush()
if len(lines) < logPageSize {
return
}
}
}
+21 -2
View File
@@ -5,6 +5,7 @@ import (
"crypto/rand"
"encoding/hex"
"encoding/json"
"errors"
"time"
"github.com/redis/go-redis/v9"
@@ -25,13 +26,31 @@ type Session struct {
var rdb *redis.Client
func InitRedis(addr string) error {
rdb = redis.NewClient(&redis.Options{Addr: addr})
// InitRedis connects the session store.
//
// Username and password may both be empty for an unauthenticated instance. For
// a legacy `requirepass` Redis, pass the password with an empty username —
// go-redis then sends AUTH with one argument instead of two.
func InitRedis(addr, username, password string) error {
rdb = redis.NewClient(&redis.Options{
Addr: addr,
Username: username,
Password: password,
})
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
return rdb.Ping(ctx).Err()
}
// PingRedis reports whether the session store is reachable. Readiness depends
// on it: a pod that cannot reach Redis can authenticate nobody.
func PingRedis(ctx context.Context) error {
if rdb == nil {
return errors.New("redis not initialised")
}
return rdb.Ping(ctx).Err()
}
func randomHex(n int) (string, error) {
b := make([]byte, n)
if _, err := rand.Read(b); err != nil {
+282
View File
@@ -0,0 +1,282 @@
// Package bus is the control plane's inter-process message bus.
//
// It exists because an agent's CommandStream lands on exactly one server
// process. With a single replica that process is also the one handling every
// REST request, so an in-memory map was enough. With several replicas it is
// not: the pod asked to run a workflow step is almost never the pod holding
// that agent's stream, and a map cannot reach across the gap.
//
// Redis is already a hard dependency (sessions), so the bus adds no new
// infrastructure. It carries three things:
//
// presence which pod, if any, currently holds an agent's stream
// commands a request/ack exchange delivering a ServerCommand to that pod
// results step results travelling back to the pod driving the run
//
// Everything here is deliberately best-effort delivery with an explicit ack
// rather than a queue. A command whose owner pod died between the presence
// check and the publish must fail loudly and immediately — the caller answers
// 503 and the operator retries — not sit in a queue waiting for a stream that
// no longer exists.
package bus
import (
"context"
"crypto/rand"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"os"
"time"
"github.com/redis/go-redis/v9"
)
// ErrNoResponder means nothing was subscribed to the channel, or the subscriber
// did not ack within the timeout. Either way the command did not reach an agent.
var ErrNoResponder = errors.New("no responder on channel")
var (
rdb *redis.Client
nodeID string
)
// Init connects the bus. It takes its own client rather than sharing the
// session store's: a subscription occupies its connection for as long as it
// lives, and every agent stream on this pod holds one, so they must not come
// out of the pool that ordinary session reads depend on.
func Init(addr, username, password string) error {
rdb = redis.NewClient(&redis.Options{
Addr: addr,
Username: username,
Password: password,
})
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
if err := rdb.Ping(ctx).Err(); err != nil {
return err
}
// Hostname is the pod name under Kubernetes, which makes a log line naming
// a node directly actionable. The random suffix keeps two processes on one
// host (or a reused pod name) from claiming each other's locks.
host, _ := os.Hostname()
if host == "" {
host = "server"
}
b := make([]byte, 4)
_, _ = rand.Read(b)
nodeID = host + "-" + hex.EncodeToString(b)
return nil
}
// NodeID identifies this process on the bus. Stable for the process lifetime.
func NodeID() string { return nodeID }
// Client exposes the underlying Redis client for the leader election and
// presence helpers in this package. It is nil before Init.
func Client() *redis.Client { return rdb }
// Channel names. Every key and channel is prefixed so a Redis shared with the
// session store (which uses km:) stays legible.
const (
prefix = "vantage:"
// CommandChannel carries envelopes to whichever pod holds serverID's stream.
CommandChannel = prefix + "cmd:"
// ackChannel carries the owner pod's answer back to the requesting pod.
ackChannel = prefix + "ack:"
// ResultChannel carries a StepResult back to the pod driving the run.
ResultChannel = prefix + "res:"
// ProxyEndChannel carries a console relay's terminal reason back to the pod
// serving the WebSocket, which is the pod that has to write the audit event.
ProxyEndChannel = prefix + "proxyend:"
// ProxyAddrChannel carries the address of a console relay listener back to
// the pod serving the WebSocket.
//
// The listener cannot be bound in advance on any particular pod. An agent's
// ProxyStream is a separate HTTP/2 request from its CommandStream, and an
// L7 proxy (Traefik) balances requests, not connections — so it may land on
// any replica, not the one holding the command stream. The pod it does land
// on binds the listener and announces it here.
ProxyAddrChannel = prefix + "proxyaddr:"
// PresenceKey records which node holds an agent's command stream.
PresenceKey = prefix + "agent:"
// ProxyPendingKey authorises one not-yet-opened ProxyStream. It is the only
// state tying a proxy_id to the instance and server it was minted for, and
// it must be visible to every replica because any of them may receive the
// stream.
ProxyPendingKey = prefix + "proxypending:"
// leaderKey records the holder of a named singleton job.
leaderKey = prefix + "leader:"
)
// Publish sends v, JSON-encoded, to channel. It reports how many subscribers
// received it, which is the only signal Redis pub/sub gives that anyone was
// listening.
func Publish(ctx context.Context, channel string, v any) (int64, error) {
b, err := json.Marshal(v)
if err != nil {
return 0, err
}
return rdb.Publish(ctx, channel, b).Result()
}
// Subscribe returns a channel of raw payloads and a function that unsubscribes.
// It blocks until Redis confirms the subscription, so a caller may publish
// immediately afterwards without racing its own subscriber.
func Subscribe(ctx context.Context, channel string) (<-chan []byte, func(), error) {
ps := rdb.Subscribe(ctx, channel)
if _, err := ps.Receive(ctx); err != nil {
_ = ps.Close()
return nil, nil, err
}
out := make(chan []byte, 64)
go func() {
defer close(out)
for msg := range ps.Channel() {
select {
case out <- []byte(msg.Payload):
default:
// A subscriber too slow to keep up would otherwise stall every
// other subscriber sharing this connection. Dropping is correct
// here: commands are acked, and a dropped ack fails the caller
// loudly rather than hanging it.
}
}
}()
return out, func() { _ = ps.Close() }, nil
}
// Request publishes req to channel and waits for a single reply on replyTo.
//
// The reply subscription is established before the publish, so a responder that
// answers instantly cannot beat the subscriber into place. A publish that
// reaches no subscriber fails immediately with ErrNoResponder rather than
// burning the whole timeout: nobody is going to answer.
func Request(ctx context.Context, channel, replyTo string, req any, timeout time.Duration) ([]byte, error) {
waitCtx, cancel := context.WithTimeout(ctx, timeout)
defer cancel()
replies, unsub, err := Subscribe(waitCtx, replyTo)
if err != nil {
return nil, fmt.Errorf("subscribe reply channel: %w", err)
}
defer unsub()
n, err := Publish(waitCtx, channel, req)
if err != nil {
return nil, err
}
if n == 0 {
return nil, ErrNoResponder
}
select {
case b, ok := <-replies:
if !ok {
return nil, ErrNoResponder
}
return b, nil
case <-waitCtx.Done():
return nil, ErrNoResponder
}
}
// Reply answers a Request on its reply channel.
func Reply(ctx context.Context, replyTo string, v any) error {
_, err := Publish(ctx, replyTo, v)
return err
}
// AckChannelFor names the reply channel for a command. The command ID is
// already unique per dispatch, so it needs no further qualification.
func AckChannelFor(commandID string) string { return ackChannel + commandID }
// SetPresence claims serverID for this node for ttl. Called repeatedly by the
// pod holding the stream; the TTL is what bounds how long a crashed pod keeps
// claiming an agent it can no longer reach.
func SetPresence(ctx context.Context, serverID string, ttl time.Duration) error {
return rdb.Set(ctx, PresenceKey+serverID, nodeID, ttl).Err()
}
// ClearPresence releases serverID, but only if this node still holds it. A
// blind DEL would let a pod whose stream had already been re-established
// elsewhere delete the new owner's claim on its way out.
func ClearPresence(ctx context.Context, serverID string) error {
return releaseIfOwner.Run(ctx, rdb, []string{PresenceKey + serverID}, nodeID).Err()
}
// PresenceHolder returns the node holding serverID's stream, or "" if none does.
func PresenceHolder(ctx context.Context, serverID string) string {
v, err := rdb.Get(ctx, PresenceKey+serverID).Result()
if err != nil {
return ""
}
return v
}
// IsConnected reports whether any pod currently holds serverID's stream.
func IsConnected(ctx context.Context, serverID string) bool {
n, err := rdb.Exists(ctx, PresenceKey+serverID).Result()
return err == nil && n > 0
}
// SetPendingProxy records that proxyID has been minted for instanceID and
// serverID, for ttl. Written before the OpenProxyCmd is dispatched, so it is in
// place before any agent can act on it.
func SetPendingProxy(ctx context.Context, proxyID, instanceID, serverID string, ttl time.Duration) error {
b, err := json.Marshal(map[string]string{"instance_id": instanceID, "server_id": serverID})
if err != nil {
return err
}
return rdb.Set(ctx, ProxyPendingKey+proxyID, b, ttl).Err()
}
// ClaimPendingProxy consumes proxyID's pending record and returns the instance
// and server it was minted for. Get and delete are one Lua call rather than two
// round trips: single use is the whole security property, and two agents
// racing the same proxy_id must not both be served.
//
// A missing record is reported as "", "" rather than an error — an unknown
// proxy_id, an expired one and a second claim are all the same refusal.
func ClaimPendingProxy(ctx context.Context, proxyID string) (instanceID, serverID string) {
v, err := claimPending.Run(ctx, rdb, []string{ProxyPendingKey + proxyID}).Text()
if err != nil || v == "" {
return "", ""
}
var rec struct {
InstanceID string `json:"instance_id"`
ServerID string `json:"server_id"`
}
if err := json.Unmarshal([]byte(v), &rec); err != nil {
return "", ""
}
return rec.InstanceID, rec.ServerID
}
// ClearPendingProxy drops a pending record whose command never reached an
// agent, so a dead proxy_id is not left claimable for the rest of its TTL.
func ClearPendingProxy(ctx context.Context, proxyID string) {
_ = rdb.Del(ctx, ProxyPendingKey+proxyID).Err()
}
var claimPending = redis.NewScript(`
local v = redis.call("GET", KEYS[1])
if v then
redis.call("DEL", KEYS[1])
return v
end
return ""
`)
var releaseIfOwner = redis.NewScript(`
if redis.call("GET", KEYS[1]) == ARGV[1] then
return redis.call("DEL", KEYS[1])
end
return 0
`)
+100
View File
@@ -0,0 +1,100 @@
package bus
import (
"context"
"log"
"time"
"github.com/redis/go-redis/v9"
)
const (
leaderTTL = 30 * time.Second
leaderRenew = 10 * time.Second
)
// RunAsLeader runs job on exactly one process at a time, cluster-wide.
//
// The background work this guards is not merely wasteful when duplicated. Every
// replica running the monitor scheduler means every check fires N times, every
// incident notification is delivered N times to the customer, and every hourly
// rollup is written N times. The Free-instance reaper is worse: it deletes
// whole instances, and two processes deleting the same one concurrently is not
// a race anyone wins.
//
// Redis rather than a Kubernetes Lease so that Docker Compose, which has no
// API server, takes the identical code path — one implementation to reason
// about, not two.
//
// job is given a context cancelled the moment leadership is lost, and must
// return when it is cancelled. Losing the lock (a paused process, a Redis
// blip) is treated as fatal to that run of the job: the successor may already
// have started, and two schedulers overlapping is the exact thing being
// prevented.
func RunAsLeader(ctx context.Context, name string, job func(context.Context)) {
go func() {
key := leaderKey + name
for {
if ctx.Err() != nil {
return
}
ok, err := rdb.SetNX(ctx, key, nodeID, leaderTTL).Result()
if err != nil {
log.Printf("leader %s: acquire failed: %v", name, err)
}
if !ok {
select {
case <-ctx.Done():
return
case <-time.After(leaderRenew):
}
continue
}
log.Printf("leader %s: acquired by %s", name, nodeID)
jobCtx, cancel := context.WithCancel(ctx)
go job(jobCtx)
holdLeadership(ctx, key, name)
cancel()
log.Printf("leader %s: released by %s", name, nodeID)
// Give up the key on a clean shutdown so a successor takes over in
// milliseconds rather than waiting out the TTL.
if ctx.Err() != nil {
relCtx, relCancel := context.WithTimeout(context.Background(), 2*time.Second)
_ = releaseIfOwner.Run(relCtx, rdb, []string{key}, nodeID).Err()
relCancel()
return
}
}
}()
}
// holdLeadership renews the lock until it is lost or ctx ends.
func holdLeadership(ctx context.Context, key, name string) {
t := time.NewTicker(leaderRenew)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
n, err := renewIfOwner.Run(ctx, rdb, []string{key}, nodeID, int(leaderTTL/time.Millisecond)).Int()
if err != nil && err != redis.Nil {
log.Printf("leader %s: renew failed, standing down: %v", name, err)
return
}
if n == 0 {
log.Printf("leader %s: lock lost, standing down", name)
return
}
}
}
}
var renewIfOwner = redis.NewScript(`
if redis.call("GET", KEYS[1]) == ARGV[1] then
return redis.call("PEXPIRE", KEYS[1], ARGV[2])
end
return 0
`)
+96
View File
@@ -123,6 +123,32 @@ type ReportChecksResponse struct{}
type ApplyUpdatesCmd struct{}
type OpenProxyCmd struct {
ProxyId string `json:"proxy_id"`
Port uint32 `json:"port"`
}
type ProxyOpen struct {
ServerId string `json:"server_id"`
AgentToken string `json:"agent_token"`
ProxyId string `json:"proxy_id"`
}
type ProxyClose struct {
Reason string `json:"reason,omitempty"`
}
type ProxyClientMsg struct {
Open *ProxyOpen `json:"open,omitempty"`
Data []byte `json:"data,omitempty"`
Close *ProxyClose `json:"close,omitempty"`
}
type ProxyServerMsg struct {
Data []byte `json:"data,omitempty"`
Close *ProxyClose `json:"close,omitempty"`
}
type ServerCommand struct {
CommandId string `json:"command_id"`
GenerateKey *GenerateKeyCmd `json:"generate_key,omitempty"`
@@ -131,6 +157,7 @@ type ServerCommand struct {
ApplyUpdates *ApplyUpdatesCmd `json:"apply_updates,omitempty"`
RunStep *RunStepCmd `json:"run_step,omitempty"`
CleanupWorkspace *CleanupWorkspaceCmd `json:"cleanup_workspace,omitempty"`
OpenProxy *OpenProxyCmd `json:"open_proxy,omitempty"`
}
type CleanupWorkspaceCmd struct {
@@ -239,6 +266,55 @@ func (c *vantageCommandStreamClient) Recv() (*ServerCommand, error) {
return m, nil
}
type Vantage_ProxyStreamServer interface {
Send(*ProxyServerMsg) error
Recv() (*ProxyClientMsg, error)
grpc.ServerStream
}
type vantageProxyStreamServer struct {
grpc.ServerStream
}
func (s *vantageProxyStreamServer) Send(m *ProxyServerMsg) error {
return s.ServerStream.SendMsg(m)
}
func (s *vantageProxyStreamServer) Recv() (*ProxyClientMsg, error) {
m := new(ProxyClientMsg)
if err := s.ServerStream.RecvMsg(m); err != nil {
return nil, err
}
return m, nil
}
type Vantage_ProxyStreamClient interface {
Send(*ProxyClientMsg) error
Recv() (*ProxyServerMsg, error)
CloseSend() error
grpc.ClientStream
}
type vantageProxyStreamClient struct {
grpc.ClientStream
}
func (c *vantageProxyStreamClient) Send(m *ProxyClientMsg) error {
return c.ClientStream.SendMsg(m)
}
func (c *vantageProxyStreamClient) Recv() (*ProxyServerMsg, error) {
m := new(ProxyServerMsg)
if err := c.ClientStream.RecvMsg(m); err != nil {
return nil, err
}
return m, nil
}
func _Vantage_ProxyStream_Handler(srv interface{}, stream grpc.ServerStream) error {
return srv.(VantageServer).ProxyStream(&vantageProxyStreamServer{stream})
}
type VantageServer interface {
Register(context.Context, *RegisterRequest) (*RegisterResponse, error)
SyncKeys(context.Context, *SyncRequest) (*SyncResponse, error)
@@ -248,6 +324,7 @@ type VantageServer interface {
SyncMonitors(context.Context, *SyncMonitorsRequest) (*SyncMonitorsResponse, error)
ReportChecks(context.Context, *ReportChecksRequest) (*ReportChecksResponse, error)
CommandStream(Vantage_CommandStreamServer) error
ProxyStream(Vantage_ProxyStreamServer) error
}
type UnimplementedVantageServer struct{}
@@ -284,6 +361,10 @@ func (UnimplementedVantageServer) CommandStream(Vantage_CommandStreamServer) err
return status.Errorf(codes.Unimplemented, "method CommandStream not implemented")
}
func (UnimplementedVantageServer) ProxyStream(Vantage_ProxyStreamServer) error {
return status.Errorf(codes.Unimplemented, "method ProxyStream not implemented")
}
type VantageClient interface {
Register(ctx context.Context, in *RegisterRequest, opts ...grpc.CallOption) (*RegisterResponse, error)
SyncKeys(ctx context.Context, in *SyncRequest, opts ...grpc.CallOption) (*SyncResponse, error)
@@ -293,6 +374,7 @@ type VantageClient interface {
SyncMonitors(ctx context.Context, in *SyncMonitorsRequest, opts ...grpc.CallOption) (*SyncMonitorsResponse, error)
ReportChecks(ctx context.Context, in *ReportChecksRequest, opts ...grpc.CallOption) (*ReportChecksResponse, error)
CommandStream(ctx context.Context, opts ...grpc.CallOption) (Vantage_CommandStreamClient, error)
ProxyStream(ctx context.Context, opts ...grpc.CallOption) (Vantage_ProxyStreamClient, error)
}
type keyManagerClient struct {
@@ -367,6 +449,14 @@ func (c *keyManagerClient) CommandStream(ctx context.Context, opts ...grpc.CallO
return &vantageCommandStreamClient{stream}, nil
}
func (c *keyManagerClient) ProxyStream(ctx context.Context, opts ...grpc.CallOption) (Vantage_ProxyStreamClient, error) {
stream, err := c.cc.NewStream(ctx, &Vantage_ServiceDesc.Streams[1], "/vantage.v1.Vantage/ProxyStream", opts...)
if err != nil {
return nil, err
}
return &vantageProxyStreamClient{stream}, nil
}
func RegisterVantageServer(s grpc.ServiceRegistrar, srv VantageServer) {
s.RegisterService(&Vantage_ServiceDesc, srv)
}
@@ -390,6 +480,12 @@ var Vantage_ServiceDesc = grpc.ServiceDesc{
ServerStreams: true,
ClientStreams: true,
},
{
StreamName: "ProxyStream",
Handler: _Vantage_ProxyStream_Handler,
ServerStreams: true,
ClientStreams: true,
},
},
Metadata: "vantage/v1/vantage.proto",
}
@@ -0,0 +1,39 @@
package pb
import (
"encoding/json"
"testing"
)
func TestProxyClientMsgRoundTrip(t *testing.T) {
in := &ProxyClientMsg{Data: []byte{0x00, 0xff, 0x10}}
raw, err := json.Marshal(in)
if err != nil {
t.Fatalf("marshal: %v", err)
}
var out ProxyClientMsg
if err := json.Unmarshal(raw, &out); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if string(out.Data) != string(in.Data) {
t.Fatalf("data mismatch: got %v want %v", out.Data, in.Data)
}
if out.Open != nil || out.Close != nil {
t.Fatalf("empty oneof fields should stay nil, got open=%v close=%v", out.Open, out.Close)
}
}
func TestOpenProxyCmdOnServerCommand(t *testing.T) {
cmd := &ServerCommand{CommandId: "c1", OpenProxy: &OpenProxyCmd{ProxyId: "p1", Port: 22}}
raw, err := json.Marshal(cmd)
if err != nil {
t.Fatalf("marshal: %v", err)
}
var out ServerCommand
if err := json.Unmarshal(raw, &out); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if out.OpenProxy == nil || out.OpenProxy.ProxyId != "p1" || out.OpenProxy.Port != 22 {
t.Fatalf("open_proxy did not round-trip: %+v", out.OpenProxy)
}
}
+56
View File
@@ -0,0 +1,56 @@
package grpcserver
import (
"log"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/grpc/pb"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/proxy"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/services"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
)
// ProxyStream carries one console TCP connection. The agent opens it after
// dialling its own loopback address, and authenticates with the same agent
// token as the command stream plus the single-use proxy_id it was handed.
func (s *vantageServer) ProxyStream(stream pb.Vantage_ProxyStreamServer) error {
msg, err := stream.Recv()
if err != nil {
return status.Errorf(codes.InvalidArgument, "expected initial open message: %v", err)
}
if msg.Open == nil {
return status.Error(codes.InvalidArgument, "first message must be open")
}
srv, err := services.ValidateAgentToken(msg.Open.ServerId, msg.Open.AgentToken)
if err != nil {
// Deliberately identical to the claim-failure response below: a bad
// token, an unknown proxy_id, and a proxy_id belonging to another
// instance/server must be indistinguishable to the caller.
log.Printf("proxy %s (server %s): invalid agent token", msg.Open.ProxyId, msg.Open.ServerId)
return status.Error(codes.PermissionDenied, "proxy session unavailable")
}
if err := serveProxy(msg.Open, srv.InstanceID, stream); err != nil {
// The reason is deliberately not returned to the agent: an unknown and a
// foreign proxy_id must be indistinguishable.
log.Printf("proxy %s (server %s): %v", msg.Open.ProxyId, msg.Open.ServerId, err)
return status.Error(codes.PermissionDenied, "proxy session unavailable")
}
return nil
}
// serveProxy claims the pending session, binds this pod's relay listener for
// it, and relays. Split out from the gRPC method so the authorisation matrix is
// testable without a real stream.
//
// The listener is bound here, on whichever replica the stream reached, rather
// than in advance on the pod holding the agent's command stream — those are not
// the same pod, because an L7 proxy balances HTTP/2 requests independently.
func serveProxy(open *pb.ProxyOpen, instanceID string, stream proxy.AgentStream) error {
sess, err := services.ClaimProxyStream(instanceID, open.ServerId, open.ProxyId)
if err != nil {
return err
}
return sess.Serve(stream)
}
+53 -6
View File
@@ -178,8 +178,13 @@ func (s *vantageServer) CommandStream(stream pb.Vantage_CommandStreamServer) err
log.Printf("update last seen %s: %v", srv.ServerID, err)
}
ch := services.Dispatcher.Connect(srv.ServerID)
defer services.Dispatcher.Disconnect(srv.ServerID)
// Serve claims this agent's presence on the bus and subscribes this pod to
// its command channel, so a dispatch issued by any other replica arrives
// here. The teardown releases both — an agent that reconnects to a
// different pod must not leave this one advertising a stream it no longer
// has.
ch, release := services.Dispatcher.Serve(stream.Context(), srv.ServerID)
defer release()
log.Printf("agent %s connected command stream", srv.ServerID)
defer log.Printf("agent %s disconnected command stream", srv.ServerID)
@@ -223,10 +228,19 @@ func (s *vantageServer) CommandStream(stream pb.Vantage_CommandStreamServer) err
}
}
func StartGRPC(port int) error {
// StartGRPC serves the agent API until stop is called.
//
// It returns a stop function rather than serving forever because an abrupt exit
// is not a neutral act here: every CommandStream handler holds an agent's
// presence claim, released by a deferred call that a killed process never runs.
// The claim then outlives its owner for the remainder of its 30s TTL, during
// which dispatch believes the agent is reachable, publishes to a channel with
// no subscriber, and fails as "agent offline" — a pod that has already exited
// still answering for an agent it can no longer reach.
func StartGRPC(port int) (stop func(), err error) {
lis, err := net.Listen("tcp", fmt.Sprintf(":%d", port))
if err != nil {
return fmt.Errorf("failed to listen: %w", err)
return nil, fmt.Errorf("failed to listen: %w", err)
}
s := grpc.NewServer(
@@ -243,6 +257,39 @@ func StartGRPC(port int) error {
)
pb.RegisterVantageServer(s, &vantageServer{})
log.Printf("gRPC server listening on :%d", port)
return s.Serve(lis)
go func() {
log.Printf("gRPC server listening on :%d", port)
if err := s.Serve(lis); err != nil {
log.Fatalf("gRPC server error: %v", err)
}
}()
// GracefulStop sends GOAWAY and waits for the handlers to return, which is
// what runs those deferred releases and, on the agent's side, ends the
// stream with a clean error it reconnects from immediately rather than
// waiting out a TCP timeout.
//
// It is bounded: an idle CommandStream returns as soon as its context is
// cancelled, but a console relay mid-transfer would otherwise hold the
// process past the pod's grace period and earn a SIGKILL — which is the
// abrupt exit this exists to avoid.
return func() {
done := make(chan struct{})
go func() {
s.GracefulStop()
close(done)
}()
select {
case <-done:
log.Println("gRPC server stopped gracefully")
case <-time.After(grpcStopTimeout):
log.Printf("gRPC server did not stop within %s, forcing", grpcStopTimeout)
s.Stop()
}
}, nil
}
// How long GracefulStop is given before outstanding streams are cut. Comfortably
// inside the chart's termination grace period, so the forced stop below still
// leaves time for the HTTP server to drain.
const grpcStopTimeout = 10 * time.Second
+37
View File
@@ -0,0 +1,37 @@
// Package proxy relays console TCP traffic between guacd and a managed server's
// agent. The agent dials only its own loopback address; the port is the single
// value it takes from the control plane.
package proxy
import (
"crypto/rand"
"encoding/hex"
"errors"
)
var (
ErrNotFound = errors.New("proxy session not found")
ErrForbidden = errors.New("proxy session belongs to another server")
)
// NewID returns a 32-byte random identifier as hex. It is the only credential
// tying an incoming ProxyStream to a pending console session.
func NewID() (string, error) {
b := make([]byte, 32)
if _, err := rand.Read(b); err != nil {
return "", err
}
return hex.EncodeToString(b), nil
}
// There is deliberately no in-process registry of pending sessions here any
// more. One existed, keyed by proxy_id, on the assumption that the pod which
// bound a listener was the pod that would receive the matching ProxyStream.
// That assumption holds only for a single replica: a ProxyStream is its own
// HTTP/2 request and an L7 proxy routes it independently of the agent's
// command stream, so with N replicas the lookup missed (N-1)/N of the time and
// the console failed with "proxy session not found".
//
// The pending record lives in Redis instead (bus.SetPendingProxy /
// ClaimPendingProxy), and the listener is bound by whichever pod the stream
// actually reaches — see services.ClaimProxyStream.
+273
View File
@@ -0,0 +1,273 @@
package proxy
import (
"errors"
"fmt"
"io"
"log"
"net"
"os"
"sync"
"time"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/grpc/pb"
)
const (
chunkSize = 32 * 1024
rendezvousTimeout = 10 * time.Second
)
// AgentStream is the server's half of a ProxyStream. It is an interface so the
// relay can be tested without gRPC.
type AgentStream interface {
Send(*pb.ProxyServerMsg) error
Recv() (*pb.ProxyClientMsg, error)
}
// Session owns one ephemeral listener and relays the single connection that
// arrives on it to an agent stream.
type Session struct {
listener net.Listener
allowed []string
timeout time.Duration
once sync.Once
mu sync.Mutex
reason string
conn net.Conn
closing bool
onEnd func(string)
rendezvous *time.Timer
}
// OnEnd registers a callback invoked exactly once, when the session finishes
// tearing down, with the recorded reason (empty if it ended cleanly).
//
// It exists because the process that observes a relay failing is not
// necessarily the process that has to record it: with several replicas the
// listener lives on the pod holding the agent's stream, while the audit event
// belongs to the pod serving the browser's WebSocket.
func (s *Session) OnEnd(fn func(string)) {
s.mu.Lock()
s.onEnd = fn
s.mu.Unlock()
}
// NewSession binds an ephemeral port on listenHost. allowed is the set of IPs
// permitted to connect; an empty set allows any, which is the degraded case
// when guacd's host could not be resolved.
//
// A watchdog is armed immediately: if the agent never opens its ProxyStream
// and calls Serve, nothing else would ever bound how long the listener (and
// the registry entry that references it) stays open. Serve cancels it once
// entered and re-arms the same deadline for the accept wait.
func NewSession(listenHost string, allowed []string) (*Session, error) {
ln, err := net.Listen("tcp", net.JoinHostPort(listenHost, "0"))
if err != nil {
return nil, fmt.Errorf("bind relay listener: %w", err)
}
s := &Session{listener: ln, allowed: allowed, timeout: rendezvousTimeout}
s.rendezvous = time.AfterFunc(rendezvousTimeout, func() {
s.Close("agent_timeout")
})
return s, nil
}
func (s *Session) Port() int {
return s.listener.Addr().(*net.TCPAddr).Port
}
// Reason reports why the session ended, empty if it ended cleanly.
func (s *Session) Reason() string {
s.mu.Lock()
defer s.mu.Unlock()
return s.reason
}
// setReason records r as the session's failure reason, first write wins. It is
// a no-op once a deliberate teardown (Close) has begun: a local Close closing
// the conn out from under the relay goroutines produces exactly the kind of
// error (net.ErrClosed, a broken pipe on write, ...) that looks like a remote
// failure but is not one, and must not overwrite — or race to set — the real
// reason, or invent one where a clean local close has none.
func (s *Session) setReason(r string) {
s.mu.Lock()
if s.reason == "" && !s.closing {
s.reason = r
}
s.mu.Unlock()
}
// Close tears the session down once. A non-empty reason is recorded only if no
// reason has been recorded already, and only before teardown begins. It closes
// both the listener and, if a connection has already been accepted, that
// connection too — an unconditional kill for the whole relay chain regardless
// of which stage it is in.
func (s *Session) Close(reason string) {
if reason != "" {
s.setReason(reason)
}
s.once.Do(func() {
// Recorded before anything is actually closed: the reader/writer
// goroutines in relay() call setReason from the errors this Close
// itself is about to cause (a closed conn, a closed stream), and
// those must be recognised as teardown noise, not a genuine failure.
s.mu.Lock()
s.closing = true
conn := s.conn
s.mu.Unlock()
if s.rendezvous != nil {
s.rendezvous.Stop()
}
_ = s.listener.Close()
if conn != nil {
_ = conn.Close()
}
// Read after teardown, not before: relay()'s goroutines set the reason
// as they unwind, and reporting one recorded earlier than that would
// name the symptom rather than the cause.
s.mu.Lock()
fn, final := s.onEnd, s.reason
s.mu.Unlock()
if fn != nil {
fn(final)
}
})
}
func (s *Session) setConn(conn net.Conn) {
s.mu.Lock()
s.conn = conn
s.mu.Unlock()
}
// Serve accepts exactly one connection, verifies its source, and relays until
// either side ends. It always closes the listener before returning.
func (s *Session) Serve(stream AgentStream) error {
defer s.Close("")
// The agent has claimed the session and opened its stream, so the
// unclaimed-rendezvous watchdog no longer applies; the accept deadline set
// just below takes over for the claimed case.
if s.rendezvous != nil {
s.rendezvous.Stop()
}
if l, ok := s.listener.(*net.TCPListener); ok {
_ = l.SetDeadline(time.Now().Add(s.timeout))
}
// Accept until a connection from an allowed source arrives or the
// rendezvous deadline (set once, above) expires. A rejected connection is
// closed and the loop retries within the same deadline; the listener is
// only closed once we have a valid connection, the deadline expires, or
// Accept fails for a genuine (non-timeout) reason.
for {
conn, err := s.listener.Accept()
if err != nil {
var netErr net.Error
if errors.Is(err, os.ErrDeadlineExceeded) || (errors.As(err, &netErr) && netErr.Timeout()) {
s.setReason("guacd_timeout")
} else {
s.setReason("accept_failed")
}
return fmt.Errorf("waiting for guacd: %w", err)
}
if !allowedRemote(conn.RemoteAddr().String(), s.allowed) {
_ = conn.Close()
continue
}
// One connection only: nothing else may claim this port. Store the
// conn first so a concurrent external Close (e.g. the browser tab
// closing) can reach it via the sync.Once teardown; then close just
// the listener directly (not through Close, which would also close
// the conn we are about to relay). The deferred s.Close("") above
// performs the real one-shot teardown, including this conn, once
// relay returns.
s.setConn(conn)
_ = s.listener.Close()
return s.relay(conn, stream)
}
}
func (s *Session) relay(conn net.Conn, stream AgentStream) error {
errCh := make(chan error, 2)
// guacd -> agent
go func() {
buf := make([]byte, chunkSize)
for {
n, err := conn.Read(buf)
if n > 0 {
chunk := make([]byte, n)
copy(chunk, buf[:n])
if sendErr := stream.Send(&pb.ProxyServerMsg{Data: chunk}); sendErr != nil {
errCh <- sendErr
return
}
}
if err != nil {
if !errors.Is(err, io.EOF) {
s.setReason("guacd_read_error")
}
errCh <- err
return
}
}
}()
// agent -> guacd
go func() {
for {
msg, err := stream.Recv()
if err != nil {
errCh <- err
return
}
if msg.Close != nil {
s.setReason(msg.Close.Reason)
errCh <- fmt.Errorf("agent closed relay: %s", msg.Close.Reason)
return
}
if len(msg.Data) > 0 {
if _, err := conn.Write(msg.Data); err != nil {
errCh <- err
return
}
}
}
}()
err := <-errCh
_ = conn.Close()
if errors.Is(err, io.EOF) {
return nil
}
return err
}
// allowedRemote reports whether remote (a host:port string) is in allowed. An
// empty allowed list permits anything.
func allowedRemote(remote string, allowed []string) bool {
host, _, err := net.SplitHostPort(remote)
if err != nil {
log.Printf("proxy: unparseable remote address %q", remote)
return false
}
if len(allowed) == 0 {
return true
}
for _, a := range allowed {
if a == host {
return true
}
}
return false
}
+11 -3
View File
@@ -15,13 +15,21 @@ import (
// An immediate pass then daily, following StartLogSweeper's shape. Daily rather
// than hourly because the unit of retention is a day: sweeping twenty-four times
// to delete the same nothing is load without a purpose.
func StartAuditSweeper() {
// It takes a context because it runs under leader election: several replicas
// all trimming the same audit logs is duplicated deletion of customer data, and
// the loop must stop the moment this process stops being the leader.
func StartAuditSweeper(ctx context.Context) {
go func() {
sweepAuditLogs()
t := time.NewTicker(24 * time.Hour)
defer t.Stop()
for range t.C {
sweepAuditLogs()
for {
select {
case <-ctx.Done():
return
case <-t.C:
sweepAuditLogs()
}
}
}()
}
+28 -11
View File
@@ -77,20 +77,37 @@ type GuacParams struct {
Params map[string]string
}
func portOr(v, def int) string {
func portOr(v, def int) int {
if v == 0 {
v = def
return def
}
return strconv.Itoa(v)
return v
}
func BuildGuacParams(srv *models.Server, protocol, sshUser, privateKey, passphrase, rdpUser, rdpPass string) (*GuacParams, error) {
host := srv.IPAddress
// TargetPort is the port on the managed server that the agent will dial on its
// own loopback address.
func TargetPort(srv *models.Server, protocol string) (int, error) {
switch protocol {
case "ssh":
return portOr(srv.SSHPort, 22), nil
case "rdp":
return portOr(srv.RDPPort, 3389), nil
case "vnc":
return 5900, nil
default:
return 0, fmt.Errorf("unsupported protocol %q", protocol)
}
}
// BuildGuacParams points guacd at the relay listener, never at the managed
// server: on a cloud deployment the server's address is not routable from here.
func BuildGuacParams(protocol, sshUser, privateKey, passphrase, rdpUser, rdpPass, relayHost string, relayPort int) (*GuacParams, error) {
port := strconv.Itoa(relayPort)
switch protocol {
case "ssh":
p := map[string]string{
"hostname": host,
"port": portOr(srv.SSHPort, 22),
"hostname": relayHost,
"port": port,
}
if sshUser == "" {
sshUser = "root"
@@ -105,8 +122,8 @@ func BuildGuacParams(srv *models.Server, protocol, sshUser, privateKey, passphra
return &GuacParams{Protocol: "ssh", Params: p}, nil
case "rdp":
return &GuacParams{Protocol: "rdp", Params: map[string]string{
"hostname": host,
"port": portOr(srv.RDPPort, 3389),
"hostname": relayHost,
"port": port,
"username": rdpUser,
"password": rdpPass,
"security": "any",
@@ -114,8 +131,8 @@ func BuildGuacParams(srv *models.Server, protocol, sshUser, privateKey, passphra
}}, nil
case "vnc":
return &GuacParams{Protocol: "vnc", Params: map[string]string{
"hostname": host,
"port": "5900",
"hostname": relayHost,
"port": port,
"password": rdpPass,
}}, nil
default:
+331
View File
@@ -0,0 +1,331 @@
package services
import (
"context"
"encoding/json"
"errors"
"fmt"
"log"
"net"
"os"
"sync"
"time"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/bus"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/grpc/pb"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/proxy"
)
// ErrAgentOffline means the console cannot be opened because the target's agent
// is not on the command stream. Every console session is relayed by the agent,
// so this is fatal rather than a degraded mode.
var ErrAgentOffline = errors.New("agent is not connected")
// A console session spans up to three processes once there is more than one
// replica, and no two of them can be assumed to be the same one:
//
// the browser's WebSocket lands on an arbitrary pod
// the agent's CommandStream lands on the pod holding presence for it
// the agent's ProxyStream lands on an arbitrary pod
//
// That third line is the one that is easy to get wrong. A ProxyStream is a
// separate HTTP/2 request, and an L7 proxy (Traefik, which the chart's gRPC
// ingress uses) balances requests rather than connections — so it does not
// follow the command stream. Binding the relay listener on the command
// stream's pod therefore fails roughly (n-1)/n of the time with "proxy session
// not found": the stream arrives at a pod whose registry is empty.
//
// So the listener is bound by whichever pod receives the ProxyStream, at the
// moment it receives it, and that pod announces its own address on
// ProxyAddrChannel. The WebSocket's pod subscribes before dispatching and
// hands the announced address to guacd. The address is the announcing pod's
// own, which is why it must resolve to a single pod (POD_IP under Kubernetes)
// rather than to the Service.
//
// Authorisation cannot live in that pod's memory either, so a pending record
// in Redis (bus.SetPendingProxy) carries the instance and server a proxy_id
// was minted for, and is consumed atomically on first claim.
//
// Teardown needs no message of its own. When the browser goes away guac closes
// its connection to the relay, the relay sees the read end, and the session
// closes itself — the same path a single-process deployment always took. Only
// the *reason* has to cross back, because the pod that writes the audit event
// is not the pod that observed the failure.
// How long Close waits for the relay's terminal reason to arrive before giving
// up. The audit event is written immediately afterwards; a reason that has not
// crossed the bus in this long is not going to improve the record by being
// waited for longer.
const proxyEndGrace = 2 * time.Second
// How long a minted proxy_id stays claimable, and how long the WebSocket's pod
// waits for the relay's address to be announced. The TTL is the longer of the
// two on purpose: a record that expired while its own opener was still waiting
// would turn a slow agent into an unexplained refusal.
const (
proxyPendingTTL = 30 * time.Second
proxyAddrWait = 15 * time.Second
)
// ConsoleProxy is a relay as seen by the pod serving the WebSocket.
type ConsoleProxy struct {
ProxyID string
Host string
Port int
serverID string
mu sync.Mutex
reason string
closed bool
stop func()
ended chan struct{}
}
// proxyEnd is the terminal event a relay's owner pod publishes.
type proxyEnd struct {
Reason string `json:"reason"`
}
func proxyListenHost() string {
if v := os.Getenv("PROXY_LISTEN_HOST"); v != "" {
return v
}
return "0.0.0.0"
}
// proxyAdvertiseHost is the address guacd will dial to reach a relay bound by
// *this* process. POD_IP wins over the configured value because with several
// replicas the configured value names the Service, and a Service cannot address
// one pod. The chart sets POD_IP from the downward API.
func proxyAdvertiseHost() string {
if v := os.Getenv("POD_IP"); v != "" {
return v
}
if v := os.Getenv("PROXY_ADVERTISE_HOST"); v != "" {
return v
}
return "server"
}
func guacdAddr() string {
if v := os.Getenv("GUACD_ADDR"); v != "" {
return v
}
return "guacd:4822"
}
// guacdHosts resolves guacd's address to the IPs allowed to claim a relay
// listener. An unresolvable host yields an empty set, which allows any source:
// refusing everything would take the console down entirely, so the narrower
// protections (ephemeral port, 10s window, single accept) carry it instead.
func guacdHosts(addr string) []string {
host, _, err := net.SplitHostPort(addr)
if err != nil {
host = addr
}
if ip := net.ParseIP(host); ip != nil {
return []string{ip.String()}
}
ips, err := net.LookupHost(host)
if err != nil {
log.Printf("proxy: cannot resolve guacd host %q, allowing any relay source: %v", host, err)
return nil
}
return ips
}
// proxyAddr is what a relay's binding pod announces: the address guacd should
// dial to reach the listener it has just bound.
type proxyAddr struct {
Host string `json:"host"`
Port int `json:"port"`
}
// ClaimProxyStream authorises an incoming ProxyStream, binds a relay listener
// for it on this pod, and announces the address to whichever pod is serving the
// browser's WebSocket. It is called from the gRPC handler, on whichever replica
// the stream happened to reach.
//
// instanceID and serverID are the *authenticated* identity of the calling
// agent; they must match the pending record or the claim is refused, so an
// agent cannot relay a console session minted for another server.
func ClaimProxyStream(instanceID, serverID, proxyID string) (*proxy.Session, error) {
ctx, cancel := context.WithTimeout(context.Background(), dispatchAckTimeout)
defer cancel()
wantInstance, wantServer := bus.ClaimPendingProxy(ctx, proxyID)
if wantInstance == "" {
return nil, proxy.ErrNotFound
}
if wantInstance != instanceID || wantServer != serverID {
return nil, proxy.ErrForbidden
}
sess, err := proxy.NewSession(proxyListenHost(), guacdHosts(guacdAddr()))
if err != nil {
return nil, err
}
sess.OnEnd(func(reason string) {
endCtx, endCancel := context.WithTimeout(context.Background(), dispatchAckTimeout)
defer endCancel()
if _, err := bus.Publish(endCtx, bus.ProxyEndChannel+proxyID, proxyEnd{Reason: reason}); err != nil {
log.Printf("proxy: publish end for %s: %v", proxyID, err)
}
})
addr := proxyAddr{Host: proxyAdvertiseHost(), Port: sess.Port()}
// Logged on the success path, not just on failure. The address guacd is
// about to be sent to is chosen per pod (POD_IP), so when a console fails
// for some replicas and not others this line is the difference between
// seeing which one answered and inferring it from silence.
log.Printf("proxy %s (server %s): relay bound on %s:%d, node %s",
proxyID, serverID, addr.Host, addr.Port, bus.NodeID())
if _, err := bus.Publish(ctx, bus.ProxyAddrChannel+proxyID, addr); err != nil {
// Nobody will ever dial this listener, so it is closed now rather than
// left to sit out its rendezvous timeout.
sess.Close("announce_failed")
return nil, fmt.Errorf("announce relay address: %w", err)
}
return sess, nil
}
// OpenConsoleProxy asks the pod holding serverID's stream to bind a relay and
// tell the agent to meet it. The caller must Close the result.
func OpenConsoleProxy(instanceID, serverID string, targetPort int) (*ConsoleProxy, error) {
if !Dispatcher.IsConnected(serverID) {
return nil, ErrAgentOffline
}
proxyID, err := proxy.NewID()
if err != nil {
return nil, fmt.Errorf("generate proxy id: %w", err)
}
// Both subscriptions are established before the command is dispatched: a
// fast agent binds and announces its relay within milliseconds, and a relay
// that fails immediately publishes its reason just as quickly. Either
// arriving before the subscriber is in place would be lost.
cp := &ConsoleProxy{ProxyID: proxyID, serverID: serverID, ended: make(chan struct{})}
endCtx, endCancel := context.WithCancel(context.Background())
ends, unsubEnd, err := bus.Subscribe(endCtx, bus.ProxyEndChannel+proxyID)
if err != nil {
endCancel()
return nil, fmt.Errorf("subscribe relay end: %w", err)
}
addrs, unsubAddr, err := bus.Subscribe(endCtx, bus.ProxyAddrChannel+proxyID)
if err != nil {
endCancel()
unsubEnd()
return nil, fmt.Errorf("subscribe relay address: %w", err)
}
cp.stop = func() {
endCancel()
unsubAddr()
unsubEnd()
}
go cp.watchEnd(ends)
// The pending record authorises the ProxyStream the agent is about to open,
// and is written before the command so it cannot lose the race with it.
pendCtx, pendCancel := context.WithTimeout(context.Background(), dispatchAckTimeout)
if err := bus.SetPendingProxy(pendCtx, proxyID, instanceID, serverID, proxyPendingTTL); err != nil {
pendCancel()
cp.stop()
return nil, fmt.Errorf("register pending relay: %w", err)
}
pendCancel()
abandon := func() {
ctx, cancel := context.WithTimeout(context.Background(), dispatchAckTimeout)
bus.ClearPendingProxy(ctx, proxyID)
cancel()
cp.stop()
}
if _, err := Dispatcher.send(CommandEnvelope{
ServerID: serverID,
Command: &pb.ServerCommand{
CommandId: proxyID,
OpenProxy: &pb.OpenProxyCmd{ProxyId: proxyID, Port: uint32(targetPort)},
},
}); err != nil {
abandon()
if errors.Is(err, ErrAgentNotConnected) {
return nil, ErrAgentOffline
}
return nil, err
}
// The command has reached the agent; the relay's address arrives only once
// the agent has actually opened its ProxyStream somewhere in the fleet.
var addr proxyAddr
select {
case b, ok := <-addrs:
if !ok {
abandon()
return nil, fmt.Errorf("relay address subscription closed")
}
if err := json.Unmarshal(b, &addr); err != nil {
abandon()
return nil, fmt.Errorf("undecodable relay address: %w", err)
}
case <-time.After(proxyAddrWait):
abandon()
return nil, fmt.Errorf("agent did not open a relay for %s", serverID)
}
if addr.Host == "" || addr.Port == 0 {
abandon()
return nil, fmt.Errorf("relay opened without an address")
}
cp.Host = addr.Host
cp.Port = addr.Port
return cp, nil
}
func (c *ConsoleProxy) watchEnd(ends <-chan []byte) {
defer close(c.ended)
b, ok := <-ends
if !ok {
return
}
var e proxyEnd
if err := json.Unmarshal(b, &e); err != nil {
return
}
c.mu.Lock()
if c.reason == "" {
c.reason = e.Reason
}
c.mu.Unlock()
}
// Close waits briefly for the relay's terminal reason, then releases the
// subscription. It is safe to call more than once.
func (c *ConsoleProxy) Close() {
c.mu.Lock()
if c.closed {
c.mu.Unlock()
return
}
c.closed = true
c.mu.Unlock()
select {
case <-c.ended:
case <-time.After(proxyEndGrace):
}
c.stop()
}
// Reason reports why the relay ended, empty if it ended cleanly.
func (c *ConsoleProxy) Reason() string {
c.mu.Lock()
defer c.mu.Unlock()
return c.reason
}
+198 -34
View File
@@ -1,63 +1,227 @@
package services
import (
"context"
"encoding/json"
"errors"
"fmt"
"log"
"net/http"
"strings"
"sync"
"time"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/bus"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/grpc/pb"
"github.com/google/uuid"
)
type commandDispatcher struct {
mu sync.RWMutex
channels map[string]chan *pb.ServerCommand
// Commands travel over the bus even when the sender is also the owner.
//
// An agent's CommandStream terminates on exactly one process, and with several
// replicas that is almost never the process handling the REST request that
// wants to talk to it. Publishing unconditionally — rather than checking for a
// local stream first and falling back — means one code path, exercised on every
// deployment including the single-replica ones, instead of a rare cross-pod
// path that only fails in production.
const (
// How long a caller waits for the owning pod to acknowledge. Generous
// enough to cross a loaded cluster, short enough that a REST handler
// answering 503 does not look hung.
dispatchAckTimeout = 5 * time.Second
// Presence must outlive a renew or two, or a momentarily slow pod would
// look offline and its agent would be declared unreachable.
presenceTTL = 30 * time.Second
presenceRenew = 10 * time.Second
)
// CommandEnvelope is what actually crosses the bus. It is the command plus the
// small amount of context the owning pod needs to act on it locally.
type CommandEnvelope struct {
ServerID string `json:"server_id"`
Command *pb.ServerCommand `json:"command"`
ReplyTo string `json:"reply_to"`
Log *LogRequest `json:"log,omitempty"`
}
var Dispatcher = &commandDispatcher{
channels: make(map[string]chan *pb.ServerCommand),
// LogRequest asks the owner pod to open a step log before it dispatches.
//
// Step output arrives on the owner pod's gRPC stream, so that is where it is
// masked and written. Shipping raw output back to the run's pod first would put
// unmasked bytes on the bus for no gain.
type LogRequest struct {
RunID string `json:"run_id"`
ServerID string `json:"server_id"`
Mask []string `json:"mask,omitempty"`
}
func (d *commandDispatcher) Connect(serverID string) chan *pb.ServerCommand {
ch := make(chan *pb.ServerCommand, 16)
d.mu.Lock()
d.channels[serverID] = ch
d.mu.Unlock()
return ch
// CommandAck is the owner pod's answer. It reports only that the command
// reached the agent's stream.
//
// A console relay listener used to be bound here and its address returned in
// this ack. It no longer is: the agent's ProxyStream does not necessarily
// arrive at the pod holding its command stream, so the listener is bound by
// whichever pod receives that stream and announced on bus.ProxyAddrChannel.
type CommandAck struct {
OK bool `json:"ok"`
Error string `json:"error,omitempty"`
Node string `json:"node,omitempty"`
}
func (d *commandDispatcher) Disconnect(serverID string) {
d.mu.Lock()
delete(d.channels, serverID)
d.mu.Unlock()
type commandDispatcher struct{}
var Dispatcher = &commandDispatcher{}
// ErrAgentNotConnected is returned when no pod holds the agent's stream.
var ErrAgentNotConnected = errors.New("agent is not connected")
// Serve subscribes this pod to serverID's command channel and claims presence
// for it, returning the channel the gRPC handler should send from and a
// teardown function.
//
// send is the local delivery: it is called on the bus goroutine and must not
// block for long, which is why it pushes onto a buffered channel rather than
// writing to the gRPC stream directly.
func (d *commandDispatcher) Serve(ctx context.Context, serverID string) (<-chan *pb.ServerCommand, func()) {
out := make(chan *pb.ServerCommand, 16)
envelopes, unsub, err := bus.Subscribe(ctx, bus.CommandChannel+serverID)
if err != nil {
// Without a subscription this pod cannot receive commands for the
// agent. Claiming presence anyway would advertise a stream nobody can
// reach, which is worse than the agent appearing offline.
log.Printf("dispatch: subscribe for %s failed, commands will not reach it: %v", serverID, err)
close(out)
return out, func() {}
}
runCtx, cancel := context.WithCancel(ctx)
if err := bus.SetPresence(runCtx, serverID, presenceTTL); err != nil {
log.Printf("dispatch: claim presence for %s: %v", serverID, err)
}
go func() {
t := time.NewTicker(presenceRenew)
defer t.Stop()
for {
select {
case <-runCtx.Done():
return
case <-t.C:
if err := bus.SetPresence(runCtx, serverID, presenceTTL); err != nil {
log.Printf("dispatch: renew presence for %s: %v", serverID, err)
}
}
}
}()
go func() {
for {
select {
case <-runCtx.Done():
return
case raw, ok := <-envelopes:
if !ok {
return
}
d.handleEnvelope(runCtx, raw, out)
}
}
}()
return out, func() {
cancel()
unsub()
relCtx, relCancel := context.WithTimeout(context.Background(), 2*time.Second)
if err := bus.ClearPresence(relCtx, serverID); err != nil {
log.Printf("dispatch: release presence for %s: %v", serverID, err)
}
relCancel()
}
}
// handleEnvelope performs the owner-pod side of a dispatch: any local setup the
// command needs, then queueing it for the stream, then the ack.
func (d *commandDispatcher) handleEnvelope(ctx context.Context, raw []byte, out chan *pb.ServerCommand) {
var env CommandEnvelope
if err := json.Unmarshal(raw, &env); err != nil {
log.Printf("dispatch: undecodable envelope: %v", err)
return
}
if env.Command == nil || env.ReplyTo == "" {
return
}
ack := CommandAck{OK: true, Node: bus.NodeID()}
if env.Log != nil {
if err := StepLogs.Open(env.Command.CommandId, env.Log.RunID, env.Log.ServerID, env.Log.Mask); err != nil {
log.Printf("dispatch: open step log for %s: %v", env.Command.CommandId, err)
}
}
select {
case out <- env.Command:
default:
ack = CommandAck{OK: false, Error: "command queue full", Node: bus.NodeID()}
}
if err := bus.Reply(ctx, env.ReplyTo, ack); err != nil {
log.Printf("dispatch: reply on %s: %v", env.ReplyTo, err)
}
}
// IsConnected reports whether any pod holds the agent's command stream.
func (d *commandDispatcher) IsConnected(serverID string) bool {
d.mu.RLock()
_, ok := d.channels[serverID]
d.mu.RUnlock()
return ok
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
defer cancel()
return bus.IsConnected(ctx, serverID)
}
func (d *commandDispatcher) dispatch(serverID string, cmd *pb.ServerCommand) error {
d.mu.RLock()
ch, ok := d.channels[serverID]
d.mu.RUnlock()
if !ok {
return fmt.Errorf("agent for server %s is not connected", serverID)
}
select {
case ch <- cmd:
return nil
default:
return fmt.Errorf("command queue full for server %s", serverID)
}
_, err := d.send(CommandEnvelope{ServerID: serverID, Command: cmd})
return err
}
func DispatchRunStep(serverID, commandID string, cmd *pb.RunStepCmd) error {
return Dispatcher.dispatch(serverID, &pb.ServerCommand{CommandId: commandID, RunStep: cmd})
// send publishes the envelope and waits for the owning pod's ack. A missing
// responder and a refusing responder are both errors: in neither case did the
// command reach the agent.
func (d *commandDispatcher) send(env CommandEnvelope) (CommandAck, error) {
if env.Command == nil || env.Command.CommandId == "" {
return CommandAck{}, fmt.Errorf("command id is required")
}
env.ReplyTo = bus.AckChannelFor(env.Command.CommandId)
ctx, cancel := context.WithTimeout(context.Background(), dispatchAckTimeout)
defer cancel()
raw, err := bus.Request(ctx, bus.CommandChannel+env.ServerID, env.ReplyTo, env, dispatchAckTimeout)
if err != nil {
if errors.Is(err, bus.ErrNoResponder) {
return CommandAck{}, fmt.Errorf("%w: %s", ErrAgentNotConnected, env.ServerID)
}
return CommandAck{}, err
}
var ack CommandAck
if err := json.Unmarshal(raw, &ack); err != nil {
return CommandAck{}, fmt.Errorf("undecodable ack: %w", err)
}
if !ack.OK {
return ack, fmt.Errorf("agent for server %s: %s", env.ServerID, ack.Error)
}
return ack, nil
}
// DispatchRunStep sends a step to the agent and asks the owning pod to capture
// its output. mask holds the secret values that must not reach the log.
func DispatchRunStep(serverID, commandID, runID string, mask []string, cmd *pb.RunStepCmd) error {
_, err := Dispatcher.send(CommandEnvelope{
ServerID: serverID,
Command: &pb.ServerCommand{CommandId: commandID, RunStep: cmd},
Log: &LogRequest{RunID: runID, ServerID: serverID, Mask: mask},
})
return err
}
func DispatchCleanupWorkspace(serverID, workspaceID string) {
+285 -88
View File
@@ -2,9 +2,8 @@ package services
import (
"bytes"
"context"
"log"
"os"
"path/filepath"
"strings"
"sync"
"time"
@@ -12,47 +11,130 @@ import (
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/db"
"go.mongodb.org/mongo-driver/v2/bson"
"go.mongodb.org/mongo-driver/v2/mongo"
"go.mongodb.org/mongo-driver/v2/mongo/options"
)
func WorkflowLogDir() string {
dir := os.Getenv("VANTAGE_WORKFLOW_LOG_DIR")
if dir == "" {
dir = filepath.Join("data", "workflow-logs")
}
_ = os.MkdirAll(dir, 0700)
return dir
}
// Workflow run logs live in MongoDB, not on disk.
//
// They used to be a file per (run, server) under a volume. That is the cheapest
// possible writer, and it is wrong the moment there is more than one server
// process: step output arrives on whichever pod holds the agent's stream, the
// run's markers are written by whichever pod started the run, and the browser
// asks for the log through whichever pod the load balancer picked. Three pods,
// one file, one local disk — two of them see an empty log.
//
// Mongo makes every pod an equal reader and writer, which is the property that
// matters. It costs writes on the hot path, so the writer batches (see
// stepLogWriter) and both caps below exist to keep a runaway step from turning
// a workflow into a database incident.
const (
// A single line longer than this is truncated. Base64 blobs and minified
// output are the usual cause; nobody reads column 9000 of a log line.
maxLogLineBytes = 8 * 1024
func ServerRunLogPath(runID, serverID string) string {
return filepath.Join(WorkflowLogDir(), runID, serverID+".log")
// A single (run, server) log stops accepting lines here, with one final
// marker saying so. 200k lines is far past what anyone reads and still
// only a few tens of MB. Without a cap, `yes` in a step fills the database.
maxLogLines = 200_000
// The writer flushes on whichever comes first. Batching is what keeps a
// chatty step to a handful of writes a second instead of one per line.
logFlushLines = 128
logFlushInterval = 250 * time.Millisecond
)
const logLinesCol = "workflow_log_lines"
const logSeqCol = "workflow_log_seq"
func logCtx() (context.Context, context.CancelFunc) {
return context.WithTimeout(context.Background(), 10*time.Second)
}
func logTS() string {
return time.Now().UTC().Format("2006-01-02T15:04:05.000") + "Z"
}
func AppendMarker(runID, serverID, text string) (int64, error) {
path := ServerRunLogPath(runID, serverID)
if err := os.MkdirAll(filepath.Dir(path), 0700); err != nil {
return 0, err
func seqID(runID, serverID string) string { return runID + "/" + serverID }
// reserveSeq atomically claims n consecutive sequence numbers for a
// (run, server) log and returns the first.
//
// The counter is a document rather than a per-process integer because two pods
// write the same log concurrently: the run's pod emits markers while the
// agent's pod emits step output. Ordering between them is only meaningful if
// they draw from the same counter.
func reserveSeq(ctx context.Context, runID, serverID string, n int) (int64, error) {
var doc struct {
Seq int64 `bson:"seq"`
}
f, err := os.OpenFile(path, os.O_APPEND|os.O_CREATE|os.O_WRONLY, 0600)
err := db.Col(logSeqCol).FindOneAndUpdate(ctx,
bson.M{"_id": seqID(runID, serverID)},
bson.M{"$inc": bson.M{"seq": int64(n)}},
options.FindOneAndUpdate().SetUpsert(true).SetReturnDocument(options.After),
).Decode(&doc)
if err != nil {
return 0, err
}
defer f.Close()
off, _ := f.Seek(0, 2)
if _, err := f.WriteString("[" + logTS() + "] " + text + "\n"); err != nil {
return off, err
}
return off, nil
return doc.Seq - int64(n), nil
}
type logLine struct {
RunID string `bson:"run_id"`
ServerID string `bson:"server_id"`
Seq int64 `bson:"seq"`
At time.Time `bson:"at"`
Line string `bson:"line"`
}
// writeLines reserves a block of sequence numbers and inserts the batch.
func writeLines(ctx context.Context, runID, serverID string, lines []string) (int64, error) {
if len(lines) == 0 {
return 0, nil
}
start, err := reserveSeq(ctx, runID, serverID, len(lines))
if err != nil {
return 0, err
}
now := time.Now().UTC()
docs := make([]any, 0, len(lines))
for i, l := range lines {
docs = append(docs, logLine{
RunID: runID,
ServerID: serverID,
Seq: start + int64(i),
At: now,
Line: l,
})
}
// Unordered: one rejected document must not discard the rest of the batch.
_, err = db.Col(logLinesCol).InsertMany(ctx, docs, options.InsertMany().SetOrdered(false))
return start, err
}
// AppendMarker writes one control line (step banners, retries, run outcome) and
// returns its sequence number, which is what a StepRun's log_offset records so
// the UI can scroll to where a step began.
func AppendMarker(runID, serverID, text string) (int64, error) {
ctx, cancel := logCtx()
defer cancel()
return writeLines(ctx, runID, serverID, []string{"[" + logTS() + "] " + text})
}
// stepLogWriter accumulates one command's output, splits it into lines, masks
// secrets and flushes batches to Mongo.
type stepLogWriter struct {
mu sync.Mutex
f *os.File
mu sync.Mutex
runID string
serverID string
secrets []string
carry []byte
secrets []string
pending []string
written int
capped bool
stop chan struct{}
done chan struct{}
}
type stepLogRegistry struct {
@@ -60,18 +142,30 @@ type stepLogRegistry struct {
writers map[string]*stepLogWriter
}
// The registry stays process-local, and correctly so: a step's output arrives
// on the pod holding that agent's stream, and that is the same pod the
// dispatch envelope asked to open the writer. Nothing here crosses pods —
// only the lines it produces do, by virtue of landing in Mongo.
var StepLogs = &stepLogRegistry{writers: make(map[string]*stepLogWriter)}
func (r *stepLogRegistry) Open(commandID, path string, secrets []string) error {
if err := os.MkdirAll(filepath.Dir(path), 0700); err != nil {
return err
func (r *stepLogRegistry) Open(commandID, runID, serverID string, secrets []string) error {
w := &stepLogWriter{
runID: runID,
serverID: serverID,
secrets: secrets,
stop: make(chan struct{}),
done: make(chan struct{}),
}
f, err := os.OpenFile(path, os.O_APPEND|os.O_CREATE|os.O_WRONLY, 0600)
if err != nil {
return err
}
w := &stepLogWriter{f: f, secrets: secrets}
go w.flushLoop()
r.mu.Lock()
if old := r.writers[commandID]; old != nil {
// A retry reuses the command ID. Close the previous attempt's writer
// rather than leaking its flush goroutine.
r.mu.Unlock()
old.close()
r.mu.Lock()
}
r.writers[commandID] = w
r.mu.Unlock()
return nil
@@ -89,24 +183,22 @@ func (r *stepLogRegistry) Append(commandID string, data []byte) {
return
}
w.mu.Lock()
defer w.mu.Unlock()
buf := append(w.carry, data...)
for {
i := bytes.IndexByte(buf, '\n')
if i < 0 {
break
}
w.writeLine(buf[:i])
w.queueLine(buf[:i])
buf = buf[i+1:]
}
w.carry = append([]byte{}, buf...)
}
full := len(w.pending) >= logFlushLines
w.mu.Unlock()
func (w *stepLogWriter) writeLine(line []byte) {
masked := maskBytes(line, w.secrets)
_, _ = w.f.WriteString("[" + logTS() + "] ")
_, _ = w.f.Write(masked)
_, _ = w.f.WriteString("\n")
if full {
w.flush()
}
}
func (r *stepLogRegistry) Close(commandID string) {
@@ -117,13 +209,72 @@ func (r *stepLogRegistry) Close(commandID string) {
if w == nil {
return
}
w.close()
}
// queueLine masks, truncates and enqueues a line. Caller holds w.mu.
func (w *stepLogWriter) queueLine(line []byte) {
if w.capped {
return
}
if w.written+len(w.pending) >= maxLogLines {
w.capped = true
w.pending = append(w.pending, "["+logTS()+"] [vantage] log truncated: this step exceeded the per-server line limit")
return
}
masked := maskBytes(line, w.secrets)
if len(masked) > maxLogLineBytes {
masked = append(masked[:maxLogLineBytes], []byte(" [truncated]")...)
}
w.pending = append(w.pending, "["+logTS()+"] "+string(masked))
}
func (w *stepLogWriter) flush() {
w.mu.Lock()
if len(w.pending) == 0 {
w.mu.Unlock()
return
}
batch := w.pending
w.pending = nil
w.written += len(batch)
runID, serverID := w.runID, w.serverID
w.mu.Unlock()
ctx, cancel := logCtx()
defer cancel()
if _, err := writeLines(ctx, runID, serverID, batch); err != nil {
// Dropped rather than retried. A log line is not worth stalling the
// step it describes, and a Mongo that cannot take writes has larger
// problems than a missing line.
log.Printf("step log: write %d line(s) for run %s: %v", len(batch), runID, err)
}
}
func (w *stepLogWriter) flushLoop() {
defer close(w.done)
t := time.NewTicker(logFlushInterval)
defer t.Stop()
for {
select {
case <-w.stop:
return
case <-t.C:
w.flush()
}
}
}
func (w *stepLogWriter) close() {
close(w.stop)
<-w.done
w.mu.Lock()
defer w.mu.Unlock()
if len(w.carry) > 0 {
w.writeLine(w.carry)
w.queueLine(w.carry)
w.carry = nil
}
_ = w.f.Close()
w.mu.Unlock()
w.flush()
}
func maskBytes(b []byte, secrets []string) []byte {
@@ -137,86 +288,132 @@ func maskBytes(b []byte, secrets []string) []byte {
return []byte(s)
}
func StartLogSweeper() {
// ReadServerRunLog returns the log for one server in a run, from sequence
// number after onwards, along with the highest sequence returned.
//
// Callers page with it rather than fetching everything: the live stream asks
// repeatedly for what is new, and the plain-text endpoint walks the whole log
// in chunks so a very large one is never held in memory whole.
func ReadServerRunLog(runID, serverID string, after int64, limit int) ([]string, int64, error) {
ctx, cancel := logCtx()
defer cancel()
cur, err := db.Col(logLinesCol).Find(ctx,
bson.M{"run_id": runID, "server_id": serverID, "seq": bson.M{"$gt": after}},
options.Find().SetSort(bson.D{{Key: "seq", Value: 1}}).SetLimit(int64(limit)),
)
if err != nil {
return nil, after, err
}
defer cur.Close(ctx)
lines := make([]string, 0, 64)
last := after
for cur.Next(ctx) {
var l logLine
if err := cur.Decode(&l); err != nil {
return lines, last, err
}
lines = append(lines, l.Line)
last = l.Seq
}
return lines, last, cur.Err()
}
// HasServerRunLog reports whether any line exists, so a log read can answer 404
// rather than 200 with an empty body.
func HasServerRunLog(runID, serverID string) bool {
ctx, cancel := logCtx()
defer cancel()
err := db.Col(logLinesCol).FindOne(ctx, bson.M{"run_id": runID, "server_id": serverID}).Err()
return err == nil
}
func StartLogSweeper(ctx context.Context) {
go func() {
sweepLogs()
t := time.NewTicker(time.Hour)
defer t.Stop()
for range t.C {
sweepLogs()
for {
select {
case <-ctx.Done():
return
case <-t.C:
sweepLogs()
}
}
}()
}
// sweepLogs deletes the lines of finished runs past their instance's retention.
//
// It walks workflow_runs rather than the log collection: retention is a
// property of the run (its instance, its finish time), and a run row is the
// only place both are recorded.
func sweepLogs() {
base := WorkflowLogDir()
entries, err := os.ReadDir(base)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute)
defer cancel()
cur, err := db.Col("workflow_runs").Find(ctx,
bson.M{"finished_at": bson.M{"$ne": nil}},
options.Find().SetProjection(bson.M{"run_id": 1, "instance_id": 1, "finished_at": 1}),
)
if err != nil {
log.Printf("log sweep: list runs: %v", err)
return
}
defer cur.Close(ctx)
cache := map[string]int{}
now := time.Now()
for _, e := range entries {
if !e.IsDir() {
continue
for cur.Next(ctx) {
var run struct {
RunID string `bson:"run_id"`
InstanceID string `bson:"instance_id"`
FinishedAt *time.Time `bson:"finished_at"`
}
runID := e.Name()
dir := filepath.Join(base, runID)
instanceID, finishedAt, found, err := runRetentionInfo(runID)
if err != nil {
log.Printf("log sweep: retention lookup failed for run %s: %v", runID, err)
continue
}
if found && finishedAt == nil {
if err := cur.Decode(&run); err != nil || run.FinishedAt == nil {
continue
}
days, ok := cache[instanceID]
days, ok := cache[run.InstanceID]
if !ok {
days = defaultRetentionDays
if instanceID != "" {
if v, err := GetWorkflowLogRetentionDays(instanceID); err == nil {
if run.InstanceID != "" {
if v, err := GetWorkflowLogRetentionDays(run.InstanceID); err == nil {
days = v
}
}
cache[instanceID] = days
cache[run.InstanceID] = days
}
if days <= 0 {
continue
}
cutoff := now.AddDate(0, 0, -days)
if found {
if finishedAt.Before(cutoff) {
_ = os.RemoveAll(dir)
}
if !run.FinishedAt.Before(now.AddDate(0, 0, -days)) {
continue
}
if fi, e := os.Stat(dir); e == nil && fi.ModTime().Before(cutoff) {
_ = os.RemoveAll(dir)
if _, err := db.Col(logLinesCol).DeleteMany(ctx, bson.M{"run_id": run.RunID}); err != nil {
log.Printf("log sweep: delete lines for run %s: %v", run.RunID, err)
continue
}
_, _ = db.Col(logSeqCol).DeleteMany(ctx, bson.M{"_id": bson.M{"$regex": "^" + run.RunID + "/"}})
}
}
const defaultRetentionDays = 30
func runRetentionInfo(runID string) (string, *time.Time, bool, error) {
ctx, cancel := wfCtx()
// EnsureLogIndexes builds the indexes the log store depends on. The compound
// index is not an optimisation: every read is a range scan over it, and without
// it a log read collection-scans every line in the database.
func EnsureLogIndexes() error {
ctx, cancel := logCtx()
defer cancel()
var run struct {
InstanceID string `bson:"instance_id"`
FinishedAt *time.Time `bson:"finished_at"`
if _, err := db.Col(logLinesCol).Indexes().CreateOne(ctx, mongo.IndexModel{
Keys: bson.D{{Key: "run_id", Value: 1}, {Key: "server_id", Value: 1}, {Key: "seq", Value: 1}},
}); err != nil {
return err
}
err := db.Col("workflow_runs").FindOne(ctx, bson.M{"run_id": runID}).Decode(&run)
if err == mongo.ErrNoDocuments {
return "", nil, false, nil
}
if err != nil {
return "", nil, false, err
}
return run.InstanceID, run.FinishedAt, true, nil
return nil
}
+61 -29
View File
@@ -1,43 +1,75 @@
package services
import (
"sync"
"context"
"encoding/json"
"log"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/bus"
"gitea.hostxtra.co.uk/mrhid6/vantage/server/internal/grpc/pb"
)
type stepResultRegistry struct {
mu sync.Mutex
pending map[string]chan *pb.StepResult
}
var StepResults = &stepResultRegistry{pending: make(map[string]chan *pb.StepResult)}
func (r *stepResultRegistry) Await(commandID string) <-chan *pb.StepResult {
ch := make(chan *pb.StepResult, 1)
r.mu.Lock()
r.pending[commandID] = ch
r.mu.Unlock()
return ch
}
func (r *stepResultRegistry) Cancel(commandID string) {
r.mu.Lock()
delete(r.pending, commandID)
r.mu.Unlock()
// Step results travel back over the bus for the same reason commands travel out
// over it: the pod driving a workflow run and the pod holding the agent's
// stream are two different processes, and a map in one of them cannot be read
// by the other.
//
// Await subscribes before the command is dispatched (see dispatchAndWait),
// which is what stops a fast agent from answering into a channel nobody is
// listening on yet.
type stepResultRegistry struct{}
var StepResults = &stepResultRegistry{}
// Await subscribes to a command's result channel. The returned cancel function
// must be called once the caller is done, whether a result arrived or not —
// it is what releases the Redis subscription.
func (r *stepResultRegistry) Await(commandID string) (<-chan *pb.StepResult, func()) {
out := make(chan *pb.StepResult, 1)
ctx, cancel := context.WithCancel(context.Background())
raw, unsub, err := bus.Subscribe(ctx, bus.ResultChannel+commandID)
if err != nil {
log.Printf("step results: subscribe for %s: %v", commandID, err)
cancel()
close(out)
return out, func() {}
}
go func() {
defer close(out)
select {
case <-ctx.Done():
return
case b, ok := <-raw:
if !ok {
return
}
var res pb.StepResult
if err := json.Unmarshal(b, &res); err != nil {
log.Printf("step results: undecodable result for %s: %v", commandID, err)
return
}
out <- &res
}
}()
return out, func() {
cancel()
unsub()
}
}
// Deliver publishes a result received from an agent. Called on the pod holding
// that agent's stream, which is not usually the pod waiting for it.
func (r *stepResultRegistry) Deliver(res *pb.StepResult) {
if res == nil {
if res == nil || res.CommandId == "" {
return
}
r.mu.Lock()
ch, ok := r.pending[res.CommandId]
if ok {
delete(r.pending, res.CommandId)
}
r.mu.Unlock()
if ok {
ch <- res
ctx, cancel := context.WithTimeout(context.Background(), dispatchAckTimeout)
defer cancel()
if _, err := bus.Publish(ctx, bus.ResultChannel+res.CommandId, res); err != nil {
log.Printf("step results: publish for %s: %v", res.CommandId, err)
}
}
+17 -10
View File
@@ -240,7 +240,6 @@ func runServer(instanceID, runID string, srvIdx int, steps []models.ResolvedStep
marker := fmt.Sprintf("===== step %d/%d: %s (%s) =====", step.Order+1, len(steps), step.Name, step.Interpreter)
offset, _ := AppendMarker(runID, serverID, marker)
logPath := ServerRunLogPath(runID, serverID)
secretsSlice := secretValues(secretVals)
commandID := uuid.New().String()
@@ -250,15 +249,17 @@ func runServer(instanceID, runID string, srvIdx int, steps []models.ResolvedStep
_, _ = AppendMarker(runID, serverID, fmt.Sprintf("retry %d/%d after failure", attempts-1, maxAttempts-1))
}
_ = StepLogs.Open(commandID, logPath, secretsSlice)
res = dispatchAndWait(serverID, commandID, &pb.RunStepCmd{
// The log writer is opened by the pod that owns this agent's
// stream, not here: that is where the output arrives, and the mask
// list travels with the dispatch so it is applied before anything
// is stored.
res = dispatchAndWait(serverID, commandID, runID, secretsSlice, &pb.RunStepCmd{
Interpreter: step.Interpreter,
Script: step.Script,
Env: cmdEnv,
TimeoutSeconds: 0,
WorkspaceId: runID,
})
StepLogs.Close(commandID)
if res != nil && res.ExitCode == 0 {
break
}
@@ -322,10 +323,14 @@ func runServer(instanceID, runID string, srvIdx int, steps []models.ResolvedStep
})
}
func dispatchAndWait(serverID, commandID string, cmd *pb.RunStepCmd) *pb.StepResult {
ch := StepResults.Await(commandID)
if err := DispatchRunStep(serverID, commandID, cmd); err != nil {
StepResults.Cancel(commandID)
// dispatchAndWait subscribes to the result before dispatching, because the two
// happen on different pods and an agent that answers quickly would otherwise
// publish into a channel this process had not yet joined.
func dispatchAndWait(serverID, commandID, runID string, mask []string, cmd *pb.RunStepCmd) *pb.StepResult {
ch, done := StepResults.Await(commandID)
defer done()
if err := DispatchRunStep(serverID, commandID, runID, mask, cmd); err != nil {
return &pb.StepResult{ExitCode: 1, Stderr: "[vantage] dispatch failed: " + err.Error()}
}
wait := time.Duration(cmd.TimeoutSeconds)*time.Second + stepDispatchGrace
@@ -333,10 +338,12 @@ func dispatchAndWait(serverID, commandID string, cmd *pb.RunStepCmd) *pb.StepRes
wait = 30*time.Minute + stepDispatchGrace
}
select {
case res := <-ch:
case res, ok := <-ch:
if !ok || res == nil {
return &pb.StepResult{ExitCode: 1, Stderr: "[vantage] lost the result channel before the agent answered"}
}
return res
case <-time.After(wait):
StepResults.Cancel(commandID)
return &pb.StepResult{ExitCode: 124, Stderr: "[vantage] timed out waiting for agent result"}
}
}
+3
View File
@@ -60,6 +60,9 @@ func EnsureWorkflowIndexes() error {
}); err != nil {
return err
}
if err := EnsureLogIndexes(); err != nil {
return err
}
_, err := db.Col("workflow_runs").Indexes().CreateOne(ctx, mongo.IndexModel{
Keys: bson.D{{Key: "run_id", Value: 1}}, Options: options.Index().SetUnique(true),
})
+120 -10
View File
@@ -6,7 +6,27 @@ import { useParams, useRouter, useSearchParams } from "next/navigation";
import Link from "next/link";
import { api } from "@/lib/api";
import { Button, Card } from "@/components/ui";
import { openConsole } from "@/lib/guacConsole";
import { openConsole, type ConsoleFailure, type ConsoleState } from "@/lib/guacConsole";
// The session's own lifecycle, which is not the same as the tunnel's: "idle"
// means the form is showing, and everything else means a session has been
// started and the viewport owns the page.
type SessionPhase = "idle" | ConsoleState;
const PHASE_LABEL: Record<Exclude<SessionPhase, "idle">, string> = {
connecting: "Connecting",
connected: "Connected",
disconnected: "Disconnected",
error: "Failed",
};
// Shape as well as colour: state must never read by colour alone.
const PHASE_DOT: Record<Exclude<SessionPhase, "idle">, string> = {
connecting: "bg-warning animate-pulse",
connected: "bg-success",
disconnected: "bg-text-tertiary",
error: "bg-danger",
};
export default function ServerConsolePage() {
const params = useParams();
@@ -29,8 +49,10 @@ export default function ServerConsolePage() {
const [rdpPassword, setRdpPassword] = useState("");
const [vncPassword, setVncPassword] = useState("");
const [connecting, setConnecting] = useState(false);
const [connected, setConnected] = useState(false);
const [phase, setPhase] = useState<SessionPhase>("idle");
const [error, setError] = useState<string | null>(null);
const [failure, setFailure] = useState<ConsoleFailure | null>(null);
const connected = phase !== "idle";
const [pending, setPending] = useState<{ token: string; wsPath: string } | null>(null);
const [zoom, setZoom] = useState(1);
const dprRef = useRef(1);
@@ -74,6 +96,7 @@ export default function ServerConsolePage() {
async function handleConnect() {
setError(null);
setFailure(null);
setConnecting(true);
try {
const body: Parameters<typeof api.connectConsole>[0] = {
@@ -93,10 +116,18 @@ export default function ServerConsolePage() {
const { token, ws_path } = await api.connectConsole(body);
// "connecting", not "connected": all we have so far is a token. The
// real state now comes from the tunnel, which is the only thing that
// knows whether the far end ever answered.
setPending({ token, wsPath: ws_path });
setConnected(true);
setPhase("connecting");
} catch (e) {
setError(e instanceof Error ? e.message : "Failed to connect");
const message = e instanceof Error ? e.message : "Failed to connect";
setError(
message.includes("agent_offline")
? "The agent on this server is not connected, so a console session cannot be opened."
: message
);
} finally {
setConnecting(false);
}
@@ -121,7 +152,10 @@ export default function ServerConsolePage() {
`&height=${Math.floor(rect.height * dpr)}` +
`&dpi=96`;
connectionRef.current = openConsole(containerRef.current, wsUrl, connectData);
connectionRef.current = openConsole(containerRef.current, wsUrl, connectData, {
onState: (s) => setPhase(s),
onFailure: (f) => setFailure(f),
});
connectionRef.current.setScale(zoom / dpr);
setPending(null);
}, [connected, pending]);
@@ -140,15 +174,23 @@ export default function ServerConsolePage() {
connectionRef.current.setScale(zoom / dpr);
}, [zoom]);
// Returns to the connection form. Used both by the Disconnect button and by
// Reconnect, which is the same teardown followed by a fresh dial.
function handleDisconnect() {
connectionRef.current?.disconnect();
connectionRef.current = null;
setConnected(false);
setPhase("idle");
setFailure(null);
if (containerRef.current) {
containerRef.current.innerHTML = "";
}
}
function handleReconnect() {
handleDisconnect();
void handleConnect();
}
if (serverLoading || keysLoading) {
return (
<div className="flex h-full items-center justify-center">
@@ -269,9 +311,18 @@ export default function ServerConsolePage() {
</Card>
) : (
<div className="mb-4 flex flex-wrap items-center gap-3">
<span className="inline-flex items-center gap-2 rounded-full border border-border bg-surface-2 px-3 py-1.5 text-xs font-medium text-text-secondary">
<span className={`h-2 w-2 rounded-full ${PHASE_DOT[phase as Exclude<SessionPhase, "idle">]}`} />
{PHASE_LABEL[phase as Exclude<SessionPhase, "idle">]}
</span>
<Button variant="danger" onClick={handleDisconnect}>
Disconnect
</Button>
{(phase === "error" || phase === "disconnected") && (
<Button variant="secondary" onClick={handleReconnect}>
Reconnect
</Button>
)}
<label className="text-sm text-text-secondary">Scale</label>
<select
value={zoom}
@@ -292,10 +343,69 @@ export default function ServerConsolePage() {
</div>
)}
<div
ref={containerRef}
className="min-h-[500px] flex-1 overflow-hidden rounded-lg border border-border bg-black"
/>
{/* The viewport is always mounted Guacamole attaches its display
element to it on connect, so it cannot be conditionally rendered.
Anything the operator needs to be told is layered over it
instead, which is what a bare black rectangle never did. */}
<div className="relative min-h-[500px] flex-1">
<div
ref={containerRef}
className="absolute inset-0 overflow-hidden rounded-lg border border-border bg-well"
/>
{phase === "connecting" && (
<div className="pointer-events-none absolute inset-0 flex flex-col items-center justify-center gap-3 rounded-lg bg-ground/70">
<div className="h-8 w-8 animate-spin rounded-full border-2 border-border border-t-accent" />
<p className="text-sm text-text-secondary">
Opening {protocol.toUpperCase()} session on {server.hostname}
</p>
</div>
)}
{(phase === "error" || phase === "disconnected") && (
<div className="absolute inset-0 flex items-center justify-center rounded-lg bg-ground/80 p-6">
<Card className="max-w-md">
<div className="space-y-3">
<div className="flex items-center gap-2">
<span
className={`h-2 w-2 rounded-full ${
phase === "error" ? "bg-danger" : "bg-text-tertiary"
}`}
/>
<h2 className="text-sm font-semibold text-text-primary">
{phase === "error" ? "Console session failed" : "Console session ended"}
</h2>
</div>
<p className="text-sm text-text-secondary">
{failure?.message ??
(phase === "error"
? "The session ended without reporting a reason."
: "The remote host closed the session.")}
</p>
{/* The numeric status is what makes a support
ticket actionable, so it is shown rather than
folded into the sentence above. */}
{failure?.code !== undefined && (
<p className="font-mono text-xs text-text-tertiary">
{failure.source === "tunnel" ? "tunnel" : "session"} status {failure.code}
</p>
)}
<div className="flex gap-2 pt-1">
<Button variant="primary" onClick={handleReconnect}>
Reconnect
</Button>
<Button variant="secondary" onClick={handleDisconnect}>
Change settings
</Button>
</div>
</div>
</Card>
</div>
)}
</div>
</div>
);
}
+15
View File
@@ -0,0 +1,15 @@
import { NextResponse } from "next/server";
// Liveness and readiness for the Next server itself. Deliberately at /healthz
// and not /api/healthz: next.config.ts rewrites the whole of /api to the Go
// server, so a probe there would report the backend's health instead of this
// process's — and would keep passing while this pod was wedged.
//
// It answers without touching the backend on purpose. web is stateless; a
// backend outage must not take every web replica out of its Service as well,
// which would turn one failure into two.
export const dynamic = "force-dynamic";
export function GET() {
return NextResponse.json({ status: "ok" });
}
+110 -21
View File
@@ -1,37 +1,135 @@
declare const Guacamole: any;
/**
* The lifecycle of a console session as the UI needs to talk about it.
*
* Deliberately not Guacamole's own state enum: "waiting" and "connecting" are
* one thing to an operator, and "disconnected because you clicked Disconnect"
* and "disconnected because the far end vanished" are two things that Guacamole
* reports identically.
*/
export type ConsoleState =
| "connecting"
| "connected"
| "disconnected"
| "error";
export type ConsoleFailure = {
/** Guacamole status code, when the failure came with one. */
code?: number;
/** Operator-facing sentence. Always set. */
message: string;
/** Where the failure was reported: the tunnel or the client session. */
source: "tunnel" | "client";
};
/**
* Guacamole status codes, as operator-facing sentences.
*
* These are the only diagnosis anyone gets from a failed console: the relay,
* guacd and the target daemon are all invisible from the browser, and guacd
* deliberately reports upstream failures as a bare number. Leaving them
* unmapped is what makes a broken console indistinguishable from a slow one.
*
* Source: Guacamole protocol status codes (guacamole-common-js).
*/
const STATUS_TEXT: Record<number, string> = {
256: "The server does not support this operation.",
512: "The remote desktop server encountered an error.",
513: "The remote desktop server is busy.",
514: "The remote host did not respond in time.",
515: "The remote host encountered an error.",
516: "The requested resource was not found.",
517: "The requested resource is already in use.",
518: "The remote connection was closed.",
519: "The remote host could not be reached.",
520: "The remote host is not currently available.",
521: "The session conflicts with another session.",
522: "The session timed out.",
523: "The session was closed.",
768: "The server rejected the connection request.",
769: "Authentication failed — check the credentials or SSH key.",
771: "Access to this connection was refused.",
776: "The session was closed after a period of inactivity.",
781: "The connection was closed because the client fell behind.",
782: "The server sent data the client could not understand.",
783: "Too many concurrent connections.",
};
/**
* describeStatus turns a Guacamole status into a sentence, always keeping the
* numeric code available separately. The code is what makes a support ticket
* actionable, so it is never discarded in favour of the prose.
*/
export function describeStatus(code: number | undefined, fallback: string): string {
if (code === undefined) return fallback;
return STATUS_TEXT[code] ?? fallback;
}
export type ConsoleHandlers = {
onState?: (state: ConsoleState) => void;
onFailure?: (failure: ConsoleFailure) => void;
};
export function openConsole(
container: HTMLElement,
wsUrl: string,
connectData = ""
connectData = "",
handlers: ConsoleHandlers = {}
): {
disconnect: () => void;
setScale: (scale: number) => void;
resize: (width: number, height: number) => void;
focus: () => void;
} {
const tunnel = new Guacamole.WebSocketTunnel(wsUrl);
const client = new Guacamole.Client(tunnel);
// Set once the caller tears the session down deliberately. Guacamole reports
// a user-initiated disconnect through exactly the same callbacks as a far-end
// failure, so without this flag closing the tab raises an error banner.
let closing = false;
const fail = (source: "tunnel" | "client", status: any, fallback: string) => {
if (closing) return;
const code: number | undefined =
typeof status?.code === "number" ? status.code : undefined;
const message = describeStatus(
code,
typeof status?.message === "string" && status.message ? status.message : fallback
);
handlers.onFailure?.({ code, message, source });
handlers.onState?.("error");
};
tunnel.onerror = (status: any) =>
fail("tunnel", status, "The connection to the control plane was lost.");
client.onerror = (status: any) =>
fail("client", status, "The remote session ended unexpectedly.");
// Guacamole.Client.State: 0 IDLE, 1 CONNECTING, 2 WAITING, 3 CONNECTED,
// 4 DISCONNECTING, 5 DISCONNECTED. CONNECTING and WAITING are one state to an
// operator — both mean "not usable yet".
client.onstatechange = (state: number) => {
if (closing) return;
if (state === 1 || state === 2) handlers.onState?.("connecting");
else if (state === 3) handlers.onState?.("connected");
else if (state === 5) handlers.onState?.("disconnected");
};
container.innerHTML = "";
container.appendChild(client.getDisplay().getElement());
container.tabIndex = 0;
container.style.outline = "none";
handlers.onState?.("connecting");
client.connect(connectData);
const display = client.getDisplay();
let scale = 1;
const mouse = new Guacamole.Mouse(display.getElement());
mouse.onmousedown = mouse.onmouseup = mouse.onmousemove = (state: any) => {
const s = new Guacamole.Mouse.State(
@@ -45,18 +143,10 @@ export function openConsole(
);
client.sendMouseState(s);
};
const keyboard = new Guacamole.Keyboard(container);
keyboard.onkeydown = (k: number) => client.sendKeyEvent(1, k);
keyboard.onkeyup = (k: number) => client.sendKeyEvent(0, k);
const refocus = () => {
if (document.activeElement !== container) container.focus({ preventScroll: true });
@@ -65,8 +155,6 @@ export function openConsole(
container.addEventListener("mousedown", refocus, true);
container.addEventListener("touchstart", refocus, true);
const onBlur = () => {
if (typeof keyboard.reset === "function") keyboard.reset();
};
@@ -77,6 +165,7 @@ export function openConsole(
return {
disconnect() {
closing = true;
container.removeEventListener("pointerdown", refocus, true);
container.removeEventListener("mousedown", refocus, true);
container.removeEventListener("touchstart", refocus, true);
File diff suppressed because one or more lines are too long