Files
ARIA 2c20b1c8a5
CI / Kotlin tests (android host + desktop) (pull_request) Successful in 8m18s
CI / Gateway plugin tests (pull_request) Failing after 15m7s
fix(http): don't let a half-open TLS connection wedge the accept loop
A client that completes TCP but vanishes mid-TLS-handshake (e.g. a
phone losing its network/VPN while traveling) blocked
ssl.SSLSocket.accept() inside serve_forever forever: the gateway
stopped accepting any new device connections (the app could not
reconnect), and on the next restart httpd.shutdown() froze the whole
event loop until the shutdown watchdog killed the process (ARIA
journal 2026-09-11 / 2026-09-23).

- Move the TLS handshake out of the accept loop: it now runs in the
  per-connection thread under a hard timeout (HANDSHAKE_TIMEOUT_S,
  10 s); a failed/timed-out handshake just closes the socket.
- stop() no longer blocks the event loop: shutdown()/server_close()/
  join run in an executor under asyncio.wait_for(10 s); if the bound
  expires the daemon threads are abandoned.
- Regression test: a silent half-open TCP connection must not stop
  fresh TLS connections from being served, and stop() must stay
  bounded.
2026-09-23 15:10:16 +02:00
..

Tests for the iris gateway plugin

Run via hermes's hermetic runner (never bare pytest)::

scripts/run_tests.sh tests/gateway/test_android.py

See docs/13-testing.md for the scenario list.

WS probe (ws_probe.py)

Manual test-client harness: connects to the real running gateway and drives a turn, printing every frame. Run with the hermes venv python (needs websockets); the gateway must already be up::

hermes-agent/.venv/bin/python tests/ws_probe.py \
    --token <IRIS_TOKEN> --send "hello"

Beyond the base modes (--send, --upload, --pull-offer, --sync, --fcm-token/--fcm-reg, --authfail, --url, --token, --device, --timeout), the probe has assertion and request modes:

  • --assert-turn — assert the turn produced message.start → ≥1 message.update → message.stop (scenario 2).
  • --assert-reasoning — assert the final message.stop carries a non-empty reasoning field (scenario 3).
  • --assert-tools — assert ≥1 tool.start with a matching tool.end (matched by index; scenario 4).
  • --assert-commentary — assert ≥1 commentary frame (scenario 5).
  • --assert-read-receipt — assert a read.receipt frame arrives after the sent message (new M7 frame; requires --send). SKIPs (exit 0, prints == SKIP: …) when the frame never arrives, e.g. against a gateway that predates the M7 frames.
  • --assert-status — assert a status frame is received (new M7 frame; SKIPs when absent).
  • --search QUERY [--scope all|chat] [--chat-id C] — send a search frame ({query, scope, limit}) and assert ≥1 hit in search.results (scenario 8). With --send, the turn is driven first, then the search.
  • --channel-create NAME / --channel-delete CHAT_ID / --channel-list — M3 channel directory management; create prints == channel created: <chat_id> for scripting.
  • --watch CHAT_ID — wait up to --timeout for a message to land in CHAT_ID (cron delivery E2E, scenario 7).
  • --offer-grace S — with --pull-offer, keep listening S seconds after the final message for a media.offer (offers are emitted post-turn, right after the final; default 15).
  • --http [--http-url http://host:port] — docs/19: drive the turn over the HTTP fallback leg instead of WS: GET /v1/health, POST /v1/frame (the message.send), receive over SSE GET /v1/events. The same assertion flags apply. The base URL defaults to the --url host with scheme ws(s) → http(s) and port 8791 (IRIS_HTTP_PORT).

Exit codes: 0 ok (incl. SKIP for absent M7 frames), 2 connect fail, 3 no hello.ack, 4 expected hello.ack, 5 authfail expected but acked, 6 timeout, 7 no final message, 8 upload/sync fail, 9 pull fail, 10 assert-turn fail, 11 assert-reasoning fail, 12 assert-tools fail, 13 assert-commentary fail, 14 search fail (error or zero hits), 15 channel.create/list fail, 16 channel.delete fail, 17 watch timeout, 18 read.receipt arrived before the sent message, 19 status frame with empty payload, 20 --http health check failed, 21 --http SSE open failed, 22 --http POST /v1/frame rejected (4xx).

E2E driver (e2e.py)

Runs the docs/13-testing.md §13.4 scenarios 1–12 automated-where- possible against the live gateway, invoking ws_probe.py (and the hermes CLI for cron) as subprocesses. Prints PASS / PARTIAL / SKIP / FAIL per scenario plus a summary table; exits 0 if no FAIL, 1 otherwise::

hermes-agent/.venv/bin/python tests/e2e.py
hermes-agent/.venv/bin/python tests/e2e.py --skip 3,5,7
hermes-agent/.venv/bin/python tests/e2e.py --url http://host:8791

The token is read from $IRIS_TOKEN, else hermes-agent/.env, else ~/.hermes/.env. The gateway must already be running (the driver never starts or stops it). It is idempotent: channels/jobs it creates are cleaned up even on failure, and leftover e2e-* channels/jobs from earlier runs are removed at start.

Scenario notes: 3 (reasoning) and 5 (commentary) are model-dependent and SKIP rather than FAIL when the current model does not emit them; 11 (push) and 12 (reconnect/sync) are PARTIAL by design — the WS leg is automated, the device-notification / gateway-kill leg is manual.