Tower 24h Log / Crash / Error Analysis

4 critical · 3 high

Window 2026-09-28 09:10 UTC → 2026-09-29 09:10 UTC · Hostinger VPS · container up 5h 00m · read-only probes

Hostinger VPS · host itself not visible from container (no docker.sock, no ssh) · disk 27G/96G 28% · load 0.05 0.12 0.09 Discord no adapter after restart 4x DNS fail · 4x 429 · 1x 404 Nous Portal auth OK · 36 API failures 429 / 500 / idle timeout hermes container · uid=10000 · PID 1 s6-svscan up 5h 00m · mem 1.5/7.8 GiB s6-svscan PID 1 · up 5h 00m Dashboard :9119 /health 302 in 2 ms 1 unreaped zombie child Gateway default PID 2726 · 2 restarts / 4 min 25 SIGTERM events in 24h Cron scheduler (in gateway) ticker heartbeat 46s · 2 active jobs last 2 runs: delivery_outcome=failed a69759413e81 10:00 IST · 5c1f24262eeb 08:30 IST state.db WAL quick_check ok · 35 MB + 4.4 MB 21 sessions · 1242 messages Profile tower-boss deleted Sep 28 12:51 (in .deleted) stale deliver target remains Traefik / Supabase / Caddy not visible from container :8080 / :3737 → no connect errors.log 24h window: 322 lines · 261 WARNING · 61 ERROR · 0 CRITICAL · 0 crashes / core dumps. Gateway restarts (SIGTERM, s6-supervise): 13 episodes / 25 events — 17:45, 17:46, 21:51 (Sep 28), 02:50, 04:09, 09:00, 09:04 UTC. Discord delivery broke at the Sep 28 17:45 restart: adapter never registers again → every cron delivery since fails. Job config says deliver discord:1542318839960965171 but the dispatcher errored on discord:1554095873460408443. tui_gateway crash-log bug recurs: SIGHUP 2026-09-29 08:17:21 UTC (also 2026-09-28 15:59:43) inside traceback.format_stack. Probes this run: /proc, ps, s6 /run/service, df, free, curl :9119, state.db (read-only), cron jobs.json + executions.db. Legend Healthy (emerald #34d399) Degraded / broken (rose #fb7185) Warning (amber #fbbf24) Not visible from the container (slate #94a3b8) Broken flow (dashed rose) Degraded / rate-limited flow (amber) Data path to state.db (violet)

Delivery pipeline

  • • Discord adapter absent after every restart
  • • 2 cron runs: completed but delivery_outcome=failed
  • • Target mismatch 1542… vs 1554…
  • • Last good delivery: Sep 28 16:26 IST

Provider health (Nous)

  • • 18x 429 "at capacity upstream"
  • • 8x 500 "gateway temporarily overloaded"
  • • 5x 429 fair-share · 2x InternalServerError
  • • No fallback_chain configured at all

Container / data plane

  • • Container up 5h 00m · PID 1 s6-svscan
  • • Load 0.05 / 0.12 / 0.09 · 2 CPU
  • • RAM 1.5 / 7.8 GiB · no swap
  • • Disk 27G / 96G (28%) · state.db ok

Critical

1 · Discord adapter never registers after a gateway restart — all cron delivery dead 7 warnings · 2 failed runs
gateway.run: No adapter available for discord fires seconds after every restart: 17:45:47, 17:46:25, 21:51:52 (Sep 28), 02:50:41, 04:09:17, 09:00:45, 09:04:14 UTC (Sep 29). executions.db confirms the consequence: job a69759413e81 (10:00 IST) and 5c1f24262eeb (08:30 IST) both completed with delivery_outcome = failed; the last successful delivery was Sep 28 16:26 IST, i.e. before the 17:45 restart. The job's deliver field is discord:1542318839960965171 but the dispatcher reported target discord:1554095873460408443, and the profile that channel belonged to (tower-boss) was deleted on Sep 28 12:51.
Fix Re-point the job to a live target and prove it: hermes cron edit a69759413e81 --deliver discord:1542318839960965171, then hermes send --to discord:1542318839960965171 --text ping. If that fails, the adapter itself is the fault — restart the gateway once, wait for "discord connected" in gateway.log before dispatch, and remove the stale tower-boss Discord credential/target from config so it cannot be resolved again.
2 · Gateway restarted twice in 4 minutes; 13 restart episodes in 24h 25 SIGTERM events
gateway-restart.log shows restarts at 09:00:34 and 09:04:04 UTC (plus 17:45, 17:46, 21:51, 02:50, 04:09), each an s6-supervise SIGTERM (parent_pid=163 s6-supervise gateway-default) followed by a fresh hermes gateway run --replace. Each restart is also what knocks Discord out (issue 1). At the 09:04 SIGTERM the shutdown diagnostic caught a hermes child of the dashboard burning 98.5% CPU.
Fix Find the caller before adding more restarts: check .hermes_history / dashboard console for the two gateway restart invocations at 09:00 and 09:04, and stop whatever issues them (likely a dashboard chat turn or a watchdog). Then add a restart backoff so s6 cannot cycle the gateway more than once per 5 minutes, and treat "adapter missing" as a failed start rather than a silent one.
3 · Notion MCP authentication is fully broken and retried on every turn 44 warnings · 10 parks · 10 OAuth errors
44x MCP OAuth setup failed for 'notion': non-interactive environment and no cached tokens, 10x MCP server 'notion' failed initial authentication, parking (with only 7 revivals), 5x OAuthRegistrationError: invalid_redirect_uri — Redirect URI must use HTTPS unless it is a loopback HTTP URI, 5x callback timeouts. The redirect-URI error means the client registration is unusable even in an interactive session, so this is not just a cron-environment problem.
Fix Either fix the OAuth client so its redirect URI is a loopback http://127.0.0.1:PORT (or an HTTPS URL), then run hermes mcp login notion from an interactive session and verify with hermes mcp list; or remove notion from the MCP server list for this profile so every turn stops paying the discovery/auth cost.
4 · Two cron jobs still pinned to dead models (OpenCode free tier / longcat-2.0:free) 9 + 4 errors · 1 open incident
4x RuntimeError: HTTP 403 … OpenCode's free tier can only be used from within OpenCode (weekly-review job bbcb93303aed, incident bbcb93_8800b6f633b1 still on record) and 9x This model is no longer free … switch to 'meituan/longcat-2.0', which surfaced to the user as a failed reply and left an undelivered obligation in state.db.
Fix Unpin the dead models: set bbcb93303aed to a working model (hermes cron edit bbcb93303aed --model upstage/solar-pro4:free) and drop longcat-2.0:free from the model chain so the fallback never lands on a paid-only model.

High

5 · Nous provider instability with no fallback configured 36 failures
18x RateLimitError: model temporarily at capacity upstream, 8x 500 gateway is temporarily overloaded, 5x fair-share rate limit, 2x InternalServerError, 3x Upstream idle timeout exceeded. Retries succeeded in most cases, but the auxiliary client logged main provider nous is unavailable and no fallback_chain / fallback_providers is configured — refusing to guess another logged-in provider.
Fix Declare a fallback so an outage degrades instead of failing: hermes fallback add (or a fallback_providers: list in config.yaml) pointing at a second provider you are already logged into.
6 · search_files tool fails against /opt/data — root-owned 0700 directory 4 failures (incl. this run)
Root cause reproduced: rg --files /opt/data exits 2 with rg: /opt/data/shared/files/logs: Permission denied (os error 13). That directory is root:root drwx------, so every content/file search over /opt/data returns "File search failed while running ripgrep" and the tool reports an error to the agent — 4 times in this window, once at the start of this very run.
Fix From the host (the agent is uid 10000 and cannot chmod it): chown hermes:hermes /opt/data/shared/files/logs && chmod 750 /opt/data/shared/files/logs. Alternatively search narrower paths, but the directory fix removes the failure permanently.
7 · Dashboard leaves an unreaped child at 98.5% CPU 1 zombie · 1 stall
PID 2701 is Zs [hermes] <defunct> with PPid 148 (the dashboard). The 09:04 shutdown diagnostic caught that same PID at 98.5% CPU. The dashboard also logged event loop stalled … (GIL pressure suspected) once in the window, consistent with a child spinning inside the dashboard process tree.
Fix Restart the dashboard service so s6 reaps the zombie (hermes gateway restart does not cover it; the dashboard is its own s6 service), and file the reaping gap: the dashboard should waitpid() console-spawned children instead of leaving them defunct.

Medium

8 · tui_gateway crash-log bug recurs on signal 2 dumps · 08:17 UTC today
SIGTERM received · 2026-09-29 08:17:21 and SIGHUP received · 2026-09-29 08:17:21 (also 2026-09-28 15:59:43) each end in entry.py line 129 _dump → traceback.format_stack failing, so the crash log is truncated exactly when it matters. The signal itself is graceful; only the dump breaks.
Fix In /opt/hermes/tui_gateway/entry.py wrap the dump(f) call in try/except Exception, and replace traceback.format_stack() with traceback.print_stack(file=f) so a linecache failure cannot clobber the log.
9 · Discord transport flakiness beyond the adapter bug 4 DNS · 4 429 · 2 send fails · 1 404
4x Cannot connect to host discord.com:443 … Temporary failure in name resolution, 4x slash-command sync rate-limited (429, backoff 80s and 152s), 2x Failed to send Discord message, and 1x 404 Not Found (10003): Unknown Channel — the last one corroborates the stale-channel theory in issue 1.
Fix Verify every configured channel/DM id still exists and belongs to a live bot (the 404 is a hard proof one does not), then reduce slash-command sync frequency to stop the 429 backoff spiral.
10 · Wrong terminal backends selected (docker + ssh) in a session 7 docker · 4 ssh
7x Docker is installed but the Docker daemon is not running / Docker backend selected but '/usr/bin/docker version' failed and 4x the SSH host and user are not configured (TERMINAL_SSH_HOST / TERMINAL_SSH_USER). Neither backend can work in this container, so each attempt burns a turn.
Fix Pin the terminal backend explicitly to local (hermes config set tools.terminal.backend local or the profile equivalent) so no session can fall through to docker/ssh.

Low

11 · Invalid toolsets configured for teams and google_chat 17 + 17 warnings
platform 'teams' has no valid toolsets configured (unknown name(s): hermes-teams) and the same for google_chat / hermes-google_chat — ~34 warnings per day with no effect other than log noise.
Fix Remove the unknown toolset names from the platform config (or the platforms themselves if unused): hermes toolsets to list valid names, then edit the platform entries.
12 · Auxiliary title generation times out 10 warnings
title_generation: request … timed out after 30.0s, then "all fallbacks exhausted" because no fallback provider is configured. Session titles silently stay unset.
Fix Raise the budget (auxiliary.title_generation.timeout) and/or point it at a small fast model rather than the main provider.
13 · Security scanner blocks legitimate cron commands 3 blocked (incl. 1 this run)
Three commands were refused mid-run: python3 -c … (script execution via -c), a grouped for … done loop ("nested executable body could not be resolved"), and earlier a cat | python3 pipe. Each one costs a wasted turn and a retry.
Fix Prefer writing a .py file and running it over python3 -c, avoid shell loops in cron commands, and if these blocks are unwanted set approvals.cron_mode: approve in config.yaml (it widens what cron may run — only do this if you accept that).
14 · Residual config noise: personality and profile leftovers 2 + 1
display.personality: mage does not match any built-in or agent.personalities entry; personality overlay will be skipped (2x), and /opt/data/profiles/ now holds only gate-guard plus a .deleted/ graveyard (tower-boss, tower-guard, watchdog…) while jobs and targets still reference tower-boss.
Fix Set a valid personality (hermes config set display.personality <name>) and purge the tower-boss references from cron job delivery targets so nothing resolves to a deleted profile.