--- id: troubleshooting title: Troubleshooting sidebar_label: Troubleshooting --- Symptoms, in the order people hit them. ## The server will not start **Exits immediately on boot.** Almost always a missing `GRPC_HOST` — the server refuses to start rather than guess a value that would break every agent later. **Fails during index creation.** The auth and settings index builders are fatal on failure by design: those unique indexes are what enforce tenant isolation, so starting without them is worse than not starting. Check the MongoDB user's permissions and whether a conflicting index already exists. **Starts, but every secret operation errors.** `KEY_ENCRYPTION_KEY` is missing or is not 64 hex characters. ## Nobody can sign in **`/setup` appears when users already exist.** The server is pointed at a different database than you think. Check the database name in `MONGO_URI` — it comes from the URI path, not a separate variable. **Sessions do not stick.** Redis is unreachable, or the cookie is being dropped because the site is served over plain HTTP. **"Wrong organisation" style rejections.** The host and session guard is comparing the request host's label against the session's organisation. Check `APP_ROOT_LABEL`. **OIDC redirects and then fails.** The callback URL registered with the provider must match exactly. Keep one local owner account so a broken provider is not a lockout. ## A server never becomes active Work through it in this order: 1. Is the agent running? `systemctl status vantage-agent`. 2. What does it say? `journalctl -u vantage-agent -f`. 3. Can that machine reach the endpoint? Test `GRPC_HOST` **from the machine**, not from the control plane host. 4. Was the token already used or expired? It is single-use and lives one hour — create a fresh enrolment rather than reusing the old command. | Symptom | Cause | | --- | --- | | Registers, then goes `offline` within minutes | Something permits the short `Register` call but drops the long-lived stream. Usually a proxy or idle-timeout middlebox | | Stays `pending` forever | Registration never happened. Token spent, or the endpoint unreachable | | Flaps between `active` and `offline` | Intermittent path, or a poll interval longer than the offline threshold | Remember the offline sweep runs every two minutes, so status is never instantaneous. ## Keys are not appearing on a machine - **It is a Windows server.** Key management is Linux-only, by design. - **The agent is not running.** Nothing polls, nothing writes. - **The key is assigned but revoked.** Revocation is soft; check the assignment state rather than the key. - **Someone edited `authorized_keys` by hand.** The agent rewrites the file to match the desired set; hand-added keys disappear on the next change. ## A workflow run fails or hangs - **Hangs at dispatch.** The target's command stream is not connected — the server may be `offline`. - **Fails immediately with an interpreter error.** A bash step on a Windows target, or PowerShell on Linux. - **A value does not reach the next step.** Values pass through the file at `$WORKFLOW_ENV`, one `KEY=value` per line. Declaring an output does not export it. - **A secret is empty.** The group is not in the step's `secret_refs`, or the key name differs from the environment variable you are reading. - **Logs stop mid-run.** A reverse proxy read timeout cut the stream. The run itself continues; reload the page. ## The console will not connect | Symptom | Cause | | --- | --- | | Connects, then closes at once | guacd unreachable. Check `GUACD_ADDR` and that the container is running | | SSH rejects the key | The stored key has no private half, or is not on the target | | RDP fails on retry | Credentials are single-use and consumed at tunnel open — enter them again | | Hangs at "connecting" | The **control plane** cannot reach the target on the protocol port. The agent's reachability is irrelevant here | | Fails only in production | The reverse proxy is not forwarding WebSocket upgrade headers | ## Monitors report down when the service is up - The check is running from the control plane and the endpoint is only reachable internally. Switch the runner to an agent on a machine that can see it. - The keyword no longer appears in the response body. - Retries are `0`, so a single dropped packet flips the state. ## Notifications are not arriving Use the channel **Test** button — it goes through the real delivery path, so a test that arrives proves credentials, network path and destination. If the test fails: a webhook returning 300 or above counts as a failure, SMTP needs `host`, `port`, `from` and `to`, and Telegram needs both `token` and `chat_id`. ## Licence problems | Symptom | Cause | | --- | --- | | `409 cloud_managed` when pasting | It is a cloud instance. Licences are written by HQ; there is nothing to paste | | Licence rejected as not matching | It is bound to a different instance UUID. Relink in HQ | | Instance degraded despite a valid-looking licence | It has expired past its grace period. Pasting still works — that endpoint stays available specifically so it can | | Cannot enrol another server | The server allowance is reached. Raise it in HQ or remove one | ## HQ portal problems **A request fails in the browser but works under `curl`.** The browser origin is missing from `ADMIN_ORIGIN`. This produces no log line at all in admin — the preflight is answered `204` without the allow-origin header, and the browser blocks the real request. **A price or plan looks wrong after an edit.** Repository variables are baked into images at build time and editing one pushes no commit, so nothing rebuilds. Trigger the build manually. See [CI/CD](../operations/ci-cd.md). ## Gathering information before asking for help ```bash docker compose ps docker compose logs --tail=200 server journalctl -u vantage-agent --no-pager -n 200 # on the affected machine ``` Include your instance UUID from **Settings → Licence** — it is the reference support works from.