--- id: troubleshooting title: Troubleshooting sidebar_label: Troubleshooting --- Symptoms, in the order people hit them. ## The server will not start **Exits immediately on boot.** Almost always a missing `GRPC_HOST`. The server refuses to start rather than guess a value that would break every agent later. **Fails while preparing the database.** Vantage stops rather than run without the safeguards it sets up at startup. Check the MongoDB user's permissions and whether an old, conflicting index is already there. **Starts, but every secret operation errors.** `KEY_ENCRYPTION_KEY` is missing or is not 64 hex characters. ## Nobody can sign in **`/setup` appears when users already exist.** The server is pointed at a different database than you think. Check the database name in `MONGO_URI`, which is taken from the end of the URI. **Sessions do not stick.** Redis is unreachable, or the cookie is being dropped because the site is served over plain HTTP. **"Wrong organisation" style rejections.** Vantage compares the address you browsed to against the instance your session belongs to. On a custom domain, check `APP_ROOT_LABEL`. **OIDC redirects and then fails.** The callback URL registered with the provider must match exactly. Keep one local owner account so a broken provider is not a lockout. ## A server never becomes active Work through it in this order: 1. Is the agent running? `systemctl status vantage-agent`. 2. What does it say? `journalctl -u vantage-agent -f`. 3. Can that machine reach the endpoint? Test `GRPC_HOST` **from the machine**, not from the control plane host. 4. Was the token already used, or older than an hour? Generate a fresh install command rather than reusing the old one. | Symptom | Cause | | --------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- | | Registers, then goes `offline` within minutes | The short registration call gets through but the long-lived connection is dropped, usually by a proxy or an idle timeout | | Stays `pending` forever | Registration never happened. Token spent, or the endpoint unreachable | | Flaps between `active` and `offline` | Intermittent path, or a poll interval longer than the offline threshold | Remember the offline sweep runs every two minutes, so status is never instantaneous. ## Keys are not appearing on a machine - **It is a Windows server.** Key management is Linux-only, by design. - **The agent is not running.** Nothing polls, nothing writes. - **The key is assigned but revoked.** Revocation is soft; check the assignment state rather than the key. - **Someone edited `authorized_keys` by hand.** The agent rewrites the file to match the desired set; hand-added keys disappear on the next change. ## A workflow run fails or hangs - **Hangs at dispatch.** The target's command stream is not connected; the server may be `offline`. - **Fails immediately with an interpreter error.** A bash step on a Windows target, or PowerShell on Linux. - **A value does not reach the next step.** Values pass through the file at `$WORKFLOW_ENV`, one `KEY=value` per line. Listing an output does not pass it on by itself. - **A secret is empty.** The group is not attached to that step, or the key name differs from the variable you are reading. - **Logs stop mid-run.** A reverse proxy read timeout cut the stream. The run itself continues; reload the page. ## The console will not connect | Symptom | Cause | | ----------------------------- | --------------------------------------------------------------------------------------------------------------- | | Connects, then closes at once | guacd unreachable. Check `GUACD_ADDR` and that the container is running | | SSH rejects the key | The stored key has no private half, or is not on the target | | RDP fails on retry | Credentials are single-use and consumed at tunnel open. Enter them again | | Hangs, then disconnects | The agent could not reach the service on that machine, or setting up the session timed out. The audit log records which | | Fails only in production | The reverse proxy is not forwarding WebSocket upgrade headers | ## Monitors report down when the service is up - The check is running from the control plane and the endpoint is only reachable internally. Switch the runner to an agent on a machine that can see it. - The keyword no longer appears in the response body. - Retries are `0`, so a single dropped packet flips the state. ## Notifications are not arriving Use the channel **Test** button. It goes through the real delivery path, so a test that arrives proves credentials, network path and destination. If the test fails: a webhook returning 300 or above counts as a failure, SMTP needs `host`, `port`, `from` and `to`, and Telegram needs both `token` and `chat_id`. ## Licence problems | Symptom | Cause | | ------------------------------------------------- | -------------------------------------------------------------------------------------------------------------- | | "Managed by Vantage HQ" when pasting | It is a cloud instance, which is licensed for you. There is nothing to paste | | Licence rejected as not matching | It was issued to a different instance ID. Relink it in Vantage HQ | | Instance degraded despite a valid-looking licence | It expired more than a few days ago. Pasting a new one still works, which is how you recover | | Cannot enrol another server | The server allowance is reached. Raise it in HQ or remove one | ## HQ portal problems The portal is a hosted service, so problems with it are ours to fix rather than yours to configure. If a page fails to load, an action reports an error, or a plan or price looks wrong after a change, contact support with your instance UUID and roughly when it happened. ## Gathering information before asking for help ```bash docker compose ps docker compose logs --tail=200 server journalctl -u vantage-agent --no-pager -n 200 # on the affected machine ``` Include your instance ID from the **Licence** page, which is the reference support works from.