Window 2026-09-28 09:10 UTC → 2026-09-29 09:10 UTC · Hostinger VPS · container up 5h 00m · read-only probes
gateway.run: No adapter available for discord fires seconds after every restart: 17:45:47, 17:46:25, 21:51:52 (Sep 28), 02:50:41, 04:09:17, 09:00:45, 09:04:14 UTC (Sep 29). executions.db confirms the consequence: job a69759413e81 (10:00 IST) and 5c1f24262eeb (08:30 IST) both completed with delivery_outcome = failed; the last successful delivery was Sep 28 16:26 IST, i.e. before the 17:45 restart. The job's deliver field is discord:1542318839960965171 but the dispatcher reported target discord:1554095873460408443, and the profile that channel belonged to (tower-boss) was deleted on Sep 28 12:51.
hermes cron edit a69759413e81 --deliver discord:1542318839960965171, then hermes send --to discord:1542318839960965171 --text ping. If that fails, the adapter itself is the fault — restart the gateway once, wait for "discord connected" in gateway.log before dispatch, and remove the stale tower-boss Discord credential/target from config so it cannot be resolved again.
gateway-restart.log shows restarts at 09:00:34 and 09:04:04 UTC (plus 17:45, 17:46, 21:51, 02:50, 04:09), each an s6-supervise SIGTERM (parent_pid=163 s6-supervise gateway-default) followed by a fresh hermes gateway run --replace. Each restart is also what knocks Discord out (issue 1). At the 09:04 SIGTERM the shutdown diagnostic caught a hermes child of the dashboard burning 98.5% CPU.
.hermes_history / dashboard console for the two gateway restart invocations at 09:00 and 09:04, and stop whatever issues them (likely a dashboard chat turn or a watchdog). Then add a restart backoff so s6 cannot cycle the gateway more than once per 5 minutes, and treat "adapter missing" as a failed start rather than a silent one.
MCP OAuth setup failed for 'notion': non-interactive environment and no cached tokens, 10x MCP server 'notion' failed initial authentication, parking (with only 7 revivals), 5x OAuthRegistrationError: invalid_redirect_uri — Redirect URI must use HTTPS unless it is a loopback HTTP URI, 5x callback timeouts. The redirect-URI error means the client registration is unusable even in an interactive session, so this is not just a cron-environment problem.
http://127.0.0.1:PORT (or an HTTPS URL), then run hermes mcp login notion from an interactive session and verify with hermes mcp list; or remove notion from the MCP server list for this profile so every turn stops paying the discovery/auth cost.
RuntimeError: HTTP 403 … OpenCode's free tier can only be used from within OpenCode (weekly-review job bbcb93303aed, incident bbcb93_8800b6f633b1 still on record) and 9x This model is no longer free … switch to 'meituan/longcat-2.0', which surfaced to the user as a failed reply and left an undelivered obligation in state.db.
bbcb93303aed to a working model (hermes cron edit bbcb93303aed --model upstage/solar-pro4:free) and drop longcat-2.0:free from the model chain so the fallback never lands on a paid-only model.
RateLimitError: model temporarily at capacity upstream, 8x 500 gateway is temporarily overloaded, 5x fair-share rate limit, 2x InternalServerError, 3x Upstream idle timeout exceeded. Retries succeeded in most cases, but the auxiliary client logged main provider nous is unavailable and no fallback_chain / fallback_providers is configured — refusing to guess another logged-in provider.
hermes fallback add (or a fallback_providers: list in config.yaml) pointing at a second provider you are already logged into.
rg --files /opt/data exits 2 with rg: /opt/data/shared/files/logs: Permission denied (os error 13). That directory is root:root drwx------, so every content/file search over /opt/data returns "File search failed while running ripgrep" and the tool reports an error to the agent — 4 times in this window, once at the start of this very run.
chown hermes:hermes /opt/data/shared/files/logs && chmod 750 /opt/data/shared/files/logs. Alternatively search narrower paths, but the directory fix removes the failure permanently.
Zs [hermes] <defunct> with PPid 148 (the dashboard). The 09:04 shutdown diagnostic caught that same PID at 98.5% CPU. The dashboard also logged event loop stalled … (GIL pressure suspected) once in the window, consistent with a child spinning inside the dashboard process tree.
hermes gateway restart does not cover it; the dashboard is its own s6 service), and file the reaping gap: the dashboard should waitpid() console-spawned children instead of leaving them defunct.
SIGTERM received · 2026-09-29 08:17:21 and SIGHUP received · 2026-09-29 08:17:21 (also 2026-09-28 15:59:43) each end in entry.py line 129 _dump → traceback.format_stack failing, so the crash log is truncated exactly when it matters. The signal itself is graceful; only the dump breaks.
/opt/hermes/tui_gateway/entry.py wrap the dump(f) call in try/except Exception, and replace traceback.format_stack() with traceback.print_stack(file=f) so a linecache failure cannot clobber the log.
Cannot connect to host discord.com:443 … Temporary failure in name resolution, 4x slash-command sync rate-limited (429, backoff 80s and 152s), 2x Failed to send Discord message, and 1x 404 Not Found (10003): Unknown Channel — the last one corroborates the stale-channel theory in issue 1.
Docker is installed but the Docker daemon is not running / Docker backend selected but '/usr/bin/docker version' failed and 4x the SSH host and user are not configured (TERMINAL_SSH_HOST / TERMINAL_SSH_USER). Neither backend can work in this container, so each attempt burns a turn.
hermes config set tools.terminal.backend local or the profile equivalent) so no session can fall through to docker/ssh.
platform 'teams' has no valid toolsets configured (unknown name(s): hermes-teams) and the same for google_chat / hermes-google_chat — ~34 warnings per day with no effect other than log noise.hermes toolsets to list valid names, then edit the platform entries.title_generation: request … timed out after 30.0s, then "all fallbacks exhausted" because no fallback provider is configured. Session titles silently stay unset.auxiliary.title_generation.timeout) and/or point it at a small fast model rather than the main provider.python3 -c … (script execution via -c), a grouped for … done loop ("nested executable body could not be resolved"), and earlier a cat | python3 pipe. Each one costs a wasted turn and a retry..py file and running it over python3 -c, avoid shell loops in cron commands, and if these blocks are unwanted set approvals.cron_mode: approve in config.yaml (it widens what cron may run — only do this if you accept that).display.personality: mage does not match any built-in or agent.personalities entry; personality overlay will be skipped (2x), and /opt/data/profiles/ now holds only gate-guard plus a .deleted/ graveyard (tower-boss, tower-guard, watchdog…) while jobs and targets still reference tower-boss.hermes config set display.personality <name>) and purge the tower-boss references from cron job delivery targets so nothing resolves to a deleted profile.