Files
vantage-docs/docs/reference/troubleshooting.md
T
mrhid6 e2c08da6a7 fix: bound and complete console relay teardown, restore proxy_failed audit
- Arm the unclaimed-relay watchdog in NewSession rather than Serve, so an
  agent that never opens its ProxyStream is bounded to 10s and reports
  reason "agent_timeout", per the design spec's failure-mode table.
- Session.Close now also closes the accepted net.Conn (stored via setConn),
  so ConsoleProxy.Close() is an unconditional kill of the whole relay chain
  instead of only closing an already-idle listener.
- Emit console.proxy_failed and end the console session from a defer in
  consoleTunnel guarded on relay.Reason(), since guac's OnDisconnect never
  runs when the connect callback errors -- which is the path every relay
  failure this feature introduces takes. Update the two docsite
  troubleshooting rows to match what the audit event can now actually show.
2026-07-31 09:21:07 +01:00

7.2 KiB

id, title, sidebar_label
id title sidebar_label
troubleshooting Troubleshooting Troubleshooting

Symptoms, in the order people hit them.

The server will not start

Exits immediately on boot. Almost always a missing GRPC_HOST the server refuses to start rather than guess a value that would break every agent later.

Fails during index creation. The auth and settings index builders are fatal on failure by design: those unique indexes are what enforce tenant isolation, so starting without them is worse than not starting. Check the MongoDB user's permissions and whether a conflicting index already exists.

Starts, but every secret operation errors. KEY_ENCRYPTION_KEY is missing or is not 64 hex characters.

Nobody can sign in

/setup appears when users already exist. The server is pointed at a different database than you think. Check the database name in MONGO_URI — it comes from the URI path, not a separate variable.

Sessions do not stick. Redis is unreachable, or the cookie is being dropped because the site is served over plain HTTP.

"Wrong organisation" style rejections. The host and session guard is comparing the request host's label against the session's organisation. Check APP_ROOT_LABEL.

OIDC redirects and then fails. The callback URL registered with the provider must match exactly. Keep one local owner account so a broken provider is not a lockout.

A server never becomes active

Work through it in this order:

  1. Is the agent running? systemctl status vantage-agent.
  2. What does it say? journalctl -u vantage-agent -f.
  3. Can that machine reach the endpoint? Test GRPC_HOST from the machine, not from the control plane host.
  4. Was the token already used or expired? It is single-use and lives one hour — create a fresh enrolment rather than reusing the old command.
Symptom Cause
Registers, then goes offline within minutes Something permits the short Register call but drops the long-lived stream. Usually a proxy or idle-timeout middlebox
Stays pending forever Registration never happened. Token spent, or the endpoint unreachable
Flaps between active and offline Intermittent path, or a poll interval longer than the offline threshold

Remember the offline sweep runs every two minutes, so status is never instantaneous.

Keys are not appearing on a machine

  • It is a Windows server. Key management is Linux-only, by design.
  • The agent is not running. Nothing polls, nothing writes.
  • The key is assigned but revoked. Revocation is soft; check the assignment state rather than the key.
  • Someone edited authorized_keys by hand. The agent rewrites the file to match the desired set; hand-added keys disappear on the next change.

A workflow run fails or hangs

  • Hangs at dispatch. The target's command stream is not connected the server may be offline.
  • Fails immediately with an interpreter error. A bash step on a Windows target, or PowerShell on Linux.
  • A value does not reach the next step. Values pass through the file at $WORKFLOW_ENV, one KEY=value per line. Declaring an output does not export it.
  • A secret is empty. The group is not in the step's secret_refs, or the key name differs from the environment variable you are reading.
  • Logs stop mid-run. A reverse proxy read timeout cut the stream. The run itself continues; reload the page.

The console will not connect

Symptom Cause
Connects, then closes at once guacd unreachable. Check GUACD_ADDR and that the container is running
SSH rejects the key The stored key has no private half, or is not on the target
RDP fails on retry Credentials are single-use and consumed at tunnel open enter them again
Hangs, then disconnects The agent never claimed the relay, nothing is listening on the protocol port on the target's own loopback address, or guacd never dialled in time. Check the audit log for console.proxy_failed — its reason (agent_timeout, dial_refused, guacd_timeout, rejected) names which
Fails only in production The reverse proxy is not forwarding WebSocket upgrade headers

Monitors report down when the service is up

  • The check is running from the control plane and the endpoint is only reachable internally. Switch the runner to an agent on a machine that can see it.
  • The keyword no longer appears in the response body.
  • Retries are 0, so a single dropped packet flips the state.

Notifications are not arriving

Use the channel Test button it goes through the real delivery path, so a test that arrives proves credentials, network path and destination.

If the test fails: a webhook returning 300 or above counts as a failure, SMTP needs host, port, from and to, and Telegram needs both token and chat_id.

Licence problems

Symptom Cause
409 cloud_managed when pasting It is a cloud instance. Licences are written by HQ; there is nothing to paste
Licence rejected as not matching It is bound to a different instance UUID. Relink in HQ
Instance degraded despite a valid-looking licence It has expired past its grace period. Pasting still works that endpoint stays available specifically so it can
Cannot enrol another server The server allowance is reached. Raise it in HQ or remove one

HQ portal problems

The portal is a hosted service, so problems with it are ours to fix rather than yours to configure. If a page fails to load, an action reports an error, or a plan or price looks wrong after a change, contact support with your instance UUID and roughly when it happened.

Gathering information before asking for help

docker compose ps
docker compose logs --tail=200 server
journalctl -u vantage-agent --no-pager -n 200   # on the affected machine

Include your instance UUID from Settings → Licence it is the reference support works from.