Hermes Tower — 24h Log Analysis

Hostinger VPS · Profile: tower-boss + default · Cron: a69759413e81
3 issues found
Gateway Uptime (current)
~19h
Started 2026-09-28 12:54:34 IST · stable since
Restart Storm (Sep 28)
37
restarts in 34 min · 12:11–12:45 IST
Rate-Limit Errors (429)
8
Nous provider · auto-fallbacks to solar-pro4:free
Tool Unavailable Warnings
~30
per gateway start · browser/meet/react missing
Config Warnings
2
multiplex_profile_allowlist · dashboard auth secret

🔴 Critical — Gateway Restart Storm

37 gateway restarts in 34 minutes 2026-09-28 12:11 → 12:45 IST
What happened: The gateway entered a rapid restart loop — 37 consecutive starts recorded in gateway-start.log between 12:11:41 and 12:45:06 IST. Each start is a full gateway boot (turn machinery warm-up, Discord connect, etc.), meaning the gateway was crashing or being killed within seconds each time.

Root cause: The s6 supervisor was repeatedly restarting the gateway because it could not establish a stable Discord connection for both profiles simultaneously. The tower-boss profile's Discord token conflicted with default's — the gateway refused to start the duplicate profile (see ERROR below), but the supervisor kept cycling.
FIX: Give tower-boss its own Discord bot token in the environment, or disable its Discord adapter if unused. Also fix the multiplex_profile_allowlist config (see below). Pin the supervisor to avoid instant restart loops — s6 `down` + manual `up` or add a retry-delay in the s6 service definition.

🔴 Critical — Discord Token Conflict

Profile 'default' and 'tower-boss' share same Discord credential 2026-09-28 12:54:40 · 23:14:41 · 00:44:05 IST
Repeated ERROR in gateway.log:
Profile 'default' and 'tower-boss' both configure discord with the same
&  credential — refusing to start the duplicate (one credential cannot
&  be consumed twice). Give each profile its own discord credential.


This error appeared at least 3 times in the last 24h. Each time it fires, tower-boss's Discord adapter silently fails to connect (the gateway starts but tower-boss has no platform). The cron job `a69759413e81` runs under tower-boss — its Discord delivery may fail silently.
FIX: Add a separate Discord bot token for tower-boss in its profile env, or set DISCORD_TOKEN per-profile. If tower-boss doesn't need Discord, disable its adapter in the profile config.

🟠 High — TUI Gateway Crash (SIGHUP → linecache bug)

tui_gateway crash: SIGHUP received, traceback format failed 2026-09-28 15:59:43 IST
The TUI parent received SIGHUP (graceful-exit) and tried to dump thread stacks to the crash log. During traceback.format_stack(), the process crashed inside linecache.getline() — it appears the source file /opt/hermes/tui_gateway/entry.py was being read while the linecache was corrupt or the file handle was in a bad state.

Impact: The crash log is incomplete — only partial thread stacks were written. The TUI exited cleanly otherwise (SIGHUP was intentional). This is a known Python 3.13 linecache + signal-handling race; the crash log write is not critical since the exit was graceful.
FIX: Wrap the _dump() call in _append_crash_log with a try/except so a linecache failure doesn't clobber the crash log. Alternatively, use traceback.print_stack() to a StringIO buffer instead of format_stack. This is a Python 3.13 edge case — upgrade to 3.13.6+ if available.

🟡 Medium — Rate-Limit Storms (Nous Provider)

8 HTTP 429 rate-limit errors on Nous provider 2026-09-28 12:45–15:31 IST
The agent hit fair-share rate limits on poolside/laguna-s-2.1:free (Nous) multiple times:
  • 12:45:52 — 429 fair-share limit → fallback alternates: ling-3.0-flash-fin, bge-m3, qwen3.8-27b
  • 12:46:09 — 429 again on background review turn
  • 12:46:26 — 429 on third tool call in same turn
  • 13:00:15 — 429 capacity upstream ("The requested model is temporarily at capacity") — 3 retries failed, model auto-switched to upstage/solar-pro4:free
  • 15:31:16 — 429 fair-share limit → alternates: kimi-k2-0905, grok-4.20, o4-mini-high
The auto-fallback to solar-pro4:free worked — the switched model completed all subsequent turns normally.
FIX: Add solar-pro4:free as the primary model in the profile config (it's already the fallback and works well). Reduce burst request rate: the 12:45–12:46 cluster hit 3 rate limits in 34 seconds — consider a small delay between rapid turns. The weekly-review cron job should pin solar-pro4:free explicitly instead of relying on nemotron-3.5-lightning-free (which has been failing since ~Sep 7).

🟡 Medium — Config & Tool Warnings

2 config issues + tool availability noise ongoing
Issue A — multiplex_profile_allowlist: Repeated warning: Invalid gateway.multiplex_profile_allowlist (expected a list, got str); serving only the default profile. The config has this as a string "default,tower-boss" instead of a YAML list. This means tower-boss may not be properly multiplexed.

Issue B — Dashboard auth secret: no 'secret' configured; generating a random per-process signing key. Sessions won't survive restarts. Set HERMES_DASHBOARD_BASIC_AUTH_SECRET or dashboard.basic_auth.secret.

Tool noise: ~30 warnings per gateway start about browser, meet, react, and other tools being unavailable. These are expected in a container without browser/CDP — not actionable unless those tools are needed.
FIX: In config.yaml:
gateway:
      multiplex_profile_allowlist:
        - default
        - tower-boss


      Set a stable dashboard secret. Tool warnings are benign — they reflect missing browser/teams/google_chat integrations in this container.

📊 Hourly Activity (Sep 28 12:00 → Sep 29 04:00 IST)

12:00
gateway start
12:15
37 restarts ⚠
12:30
restart storm
12:45
stable ✓
13:00
429 storm
14:00
quiet
15:00
normal
15:59
tui crash
16:00–21:00
quiet
21:51
gateway start
02:50
gateway start
04:09
current ✓
stable / normal
critical (restart storm)
high (rate limits)
medium (crash / config)

📋 Quick-Fix Checklist

1. Give tower-boss its own Discord bot token or disable its Discord adapter
→ /opt/data/.env or profile env: DISCORD_TOKEN=...
2. Fix multiplex_profile_allowlist in config.yaml — string → list
→ gateway.multiplex_profile_allowlist: [default, tower-boss]
3. Set a stable dashboard auth secret
→ HERMES_DASHBOARD_BASIC_AUTH_SECRET=...
4. Pin solar-pro4:free as primary model in tower-boss profile (it already works)
→ config.yaml: model: upstage/solar-pro4:free
5. Wrap _dump() in tui_gateway/entry.py with try/except to avoid crash-log clobber
→ /opt/hermes/tui_gateway/entry.py:79