Files
iris_x_hermes/gateway-plugin/tests
ARIA e6015033b6
CI / Gateway plugin tests (push) Successful in 5m9s
CI / Kotlin tests (android host + desktop) (push) Successful in 6m55s
HTTP transport: drop WS server, offline send queue + dead-stream watchdog
Gateway (docs/19):
- Remove ws_server.py; frame dispatch factored into dispatch.py
- http_server: media upload/pull, pairing over HTTP
- protocol: media frames mirrored; tests + ws_probe updated for HTTP

App:
- HttpGateway: postFrame/uploadMedia/pullMedia no longer throw on
  network failure (PostResult ok=false / Result.failure) — uncaught
  SocketTimeoutException on Dispatchers.Default crashed the app
- GatewayClient: dead-stream watchdog (health probe every 10s, 2
  failures -> redial in ~20s instead of the 45s SSE read timeout);
  state flips to Reconnecting when the stream dies, restored from the
  last hello.ack on long-poll success; poke() + backoff reset on app
  resume (MainActivity.onResume)
- Offline sends: composer enabled while disconnected; a send with no
  response (status 0) stays queued (Pending) and is auto-resent on the
  next (re)connect after a 2s outbox-replay grace; gateway 4xx
  rejections fail the bubble (tap to retry, no auto-loop)
- ChatStore: echo-replace and thread-relocate also match Failed
  bubbles (POST response lost in a network drop); loadHistory dedupes
  local failed bubbles the server already has; failMessage()
- MainActivity: poke() on resume so a backgrounded app reconnects
  promptly instead of waiting out the backoff
2026-08-22 20:10:05 +02:00
..

Tests for the android gateway plugin.

Run via hermes's hermetic runner (never bare pytest)::

scripts/run_tests.sh tests/gateway/test_android.py

See docs/13-testing.md for the scenario list.

WS probe (ws_probe.py)

Manual test-client harness: connects to the real running gateway and drives a turn, printing every frame. Run with the hermes venv python (needs websockets); the gateway must already be up::

hermes-agent/.venv/bin/python gateway-plugin/tests/ws_probe.py \
    --token <ANDROID_TOKEN> --send "hello"

Beyond the base modes (--send, --upload, --pull-offer, --sync, --fcm-token/--fcm-reg, --authfail, --url, --token, --device, --timeout), the probe has assertion and request modes:

  • --assert-turn — assert the turn produced message.start → ≥1 message.update → message.stop (scenario 2).
  • --assert-reasoning — assert the final message.stop carries a non-empty reasoning field (scenario 3).
  • --assert-tools — assert ≥1 tool.start with a matching tool.end (matched by index; scenario 4).
  • --assert-commentary — assert ≥1 commentary frame (scenario 5).
  • --assert-read-receipt — assert a read.receipt frame arrives after the sent message (new M7 frame; requires --send). SKIPs (exit 0, prints == SKIP: …) when the frame never arrives, e.g. against a gateway that predates the M7 frames.
  • --assert-status — assert a status frame is received (new M7 frame; SKIPs when absent).
  • --search QUERY [--scope all|chat] [--chat-id C] — send a search frame ({query, scope, limit}) and assert ≥1 hit in search.results (scenario 8). With --send, the turn is driven first, then the search.
  • --channel-create NAME / --channel-delete CHAT_ID / --channel-list — M3 channel directory management; create prints == channel created: <chat_id> for scripting.
  • --watch CHAT_ID — wait up to --timeout for a message to land in CHAT_ID (cron delivery E2E, scenario 7).
  • --offer-grace S — with --pull-offer, keep listening S seconds after the final message for a media.offer (offers are emitted post-turn, right after the final; default 15).
  • --http [--http-url http://host:port] — docs/19: drive the turn over the HTTP fallback leg instead of WS: GET /v1/health, POST /v1/frame (the message.send), receive over SSE GET /v1/events. The same assertion flags apply. The base URL defaults to the --url host with scheme ws(s) → http(s) and port 8791 (ANDROID_HTTP_PORT).

Exit codes: 0 ok (incl. SKIP for absent M7 frames), 2 connect fail, 3 no hello.ack, 4 expected hello.ack, 5 authfail expected but acked, 6 timeout, 7 no final message, 8 upload/sync fail, 9 pull fail, 10 assert-turn fail, 11 assert-reasoning fail, 12 assert-tools fail, 13 assert-commentary fail, 14 search fail (error or zero hits), 15 channel.create/list fail, 16 channel.delete fail, 17 watch timeout, 18 read.receipt arrived before the sent message, 19 status frame with empty payload, 20 --http health check failed, 21 --http SSE open failed, 22 --http POST /v1/frame rejected (4xx).

E2E driver (e2e.py)

Runs the docs/13-testing.md §13.4 scenarios 1–12 automated-where- possible against the live gateway, invoking ws_probe.py (and the hermes CLI for cron) as subprocesses. Prints PASS / PARTIAL / SKIP / FAIL per scenario plus a summary table; exits 0 if no FAIL, 1 otherwise::

hermes-agent/.venv/bin/python gateway-plugin/tests/e2e.py
hermes-agent/.venv/bin/python gateway-plugin/tests/e2e.py --skip 3,5,7
hermes-agent/.venv/bin/python gateway-plugin/tests/e2e.py --url ws://host:8790/ws

The token is read from $ANDROID_TOKEN, else hermes-agent/.env, else ~/.hermes/.env. The gateway must already be running (the driver never starts or stops it). It is idempotent: channels/jobs it creates are cleaned up even on failure, and leftover e2e-* channels/jobs from earlier runs are removed at start.

Scenario notes: 3 (reasoning) and 5 (commentary) are model-dependent and SKIP rather than FAIL when the current model does not emit them; 11 (push) and 12 (reconnect/sync) are PARTIAL by design — the WS leg is automated, the device-notification / gateway-kill leg is manual.