Files
iris_x_hermes/docs/19-http-fallback-transport.md
T
ARIA 2349a95dd4 HTTP fallback leg (docs/19): e2e scenario 13 + docs
- e2e.py: scenario 13 (http fallback) — drives a full turn over the
  HTTP leg (health + POST /v1/frame + SSE /v1/events, no WS) and
  asserts the user echo lands on the SSE stream in < 1.5 s.
- ws_probe.py --http: prints '== user echo in X.XXs' (the docs/19
  sendable-in-fallback timing assertion) alongside the existing
  health/POST/SSE output; same assertion flags as the WS leg.
- docs: 19 status flipped to implemented; 09-pairing-security §9.4
  cross-reference (second door, same lock: token + device allowlist,
  64 KiB cap, rate limit, optional TLS, unauthenticated /v1/health);
  13-testing manual scenario 15 + automated pointers.
2026-08-22 14:34:14 +02:00

18 KiB
Raw Blame History

19 — HTTP Fallback Transport (the "HTTP leg")

A second, short-lived-connection transport next to the WebSocket: the same JSON frames, the same outbox/cursor, the same token — served over plain HTTP by the gateway. When the WS is down (flaky network, NAT timeout, app just relaunched), the app sends over POST and receives over SSE instead of waiting 2–20 s for a WS redial.

Status: implemented (gateway leg: gateway-plugin/http_server.py; app leg: app/shared/src/commonMain/kotlin/iris/net/HttpGateway.kt + GatewayClient.State.HttpFallback). Complements — does not replace — 04-wire-protocol.md (frames), 08-push.md (outbox/sync/push), and 09-pairing-security.md (auth model).

19.1 Problem

Today the WS is the only transport, and the app hard-gates sending on a live socket (ChatScreen.doSend() no-ops unless State.Connected; GatewayClient.sendMessage() drops when socket == null). Consequences:

  • App killed → reopened: full cold dial (TCP + TLS + hello/hello.ack, 15 s dial timeout) before the user can send. On a flaky network the first dial often fails → backoff → second dial. Observed: 2–20 s of "can't send".
  • Long-lived WS is the most fragile connection type on mobile: idle sockets expire in router/CGNAT NAT tables, die on WiFi↔cellular handover, and are killed aggressively by OEM power management (MIUI on the test device). There is no foreground service holding the WS.
  • Stale detection is slow: 20 s ping interval, 60 s reap — a dead-but- unclosed socket can sit for up to a minute before redial.

19.2 Why HTTP (and why not the alternatives)

Short-lived HTTP requests are dramatically more resilient on mobile networks than a long-lived socket: no NAT table entry to expire, no proxy idle-kill, each request is a fresh connection (fast with TLS resumption), and they work through the restrictive proxies that mangle WebSockets. Sending a message becomes a single POST that completes in well under a second on a LAN — independent of whether the WS is up.

Alternatives considered and rejected (research, 2026-08):

Option Verdict
MQTT broker (QoS 1, persistent sessions) Best protocol for flaky links, but new infra (broker process) + new Python dep (paho-mqtt, breaks the zero-new-deps rule) + new Kotlin dep + frame↔topic bridge. Overkill for a 1-user agent.
ntfy as the send path (app publishes to a topic the gateway subscribes to) Adds a third party to the critical send path; public ntfy.sh is already known-flaky. Not worth it.
WebTransport / QUIC The real fix for handover flakiness (connection migration), but no OkHttp support and aioquic is a new Python dep. Future option if this doc's approach is still not enough.
gRPC New deps both sides; no advantage over WS+SSE here.
Inverted connection (app runs a local HTTP server, gateway pushes to the phone) LAN-only, breaks on cellular/remote, security mess. Rejected.
Matrix / full chat server Massive overkill for a personal agent.

Zero new Python dependencies is preserved: the HTTP leg is stdlib http.server (a ThreadingHTTPServer in a daemon thread) bridged into the gateway's asyncio loop. The app side uses the OkHttp it already depends on (hand-rolled SSE reader — the format is trivial; okhttp-eventsource is an acceptable alternative if preferred).

19.3 Shape

                        ┌──────────────────────── hermes gateway process ───────────────────────┐
                        │  AndroidAdapter                                                       │
                        │    │  frames (same protocol.Frame objects)                            │
                        │    ▼                                                                  │
                        │  _broadcast_or_log ──► outbox.append(cursor) ──► push (if no live)    │
                        │    │                │                                                 │
                        │    ▼                ▼                                                 │
                        │  WsServer (asyncio, :8790)   HttpServer (stdlib thread, :8791)        │
                        │  primary: full protocol        fallback: POST /v1/frame,              │
                        │  incl. binary media            GET /v1/events (SSE), /v1/poll         │
                        └───────────────┬──────────────────────────────┬────────────────────────┘
                                        │ WS (primary)                 │ HTTP (fallback)
                              ┌─────────┴──────────────────────────────┴────────┐
                              │  APP: transport state machine                   │
                              │  WS up   → WS only (media works, lowest latency)│
                              │  WS down → send via POST, receive via SSE/poll  │
                              └─────────────────────────────────────────────────┘

v1 scope

Over HTTP (v1) WS-only (v1)
All JSON request frames (message.send, search, channel.*, commands.catalog, agent.stop/agent.steer, …) via one generic endpoint Binary media upload (chunked binary frames)
All event/response frames via SSE (or long-poll) Binary media pull stream
sync catch-up (same outbox, same cursor) —

Media stays WS-only in v1: it is the one part of the protocol that is inherently binary/streaming, and attachments are a rarer action than sending text. While in HTTP-fallback mode the composer disables the attach button ("media needs the live connection"). HTTP media endpoints are a v2 item (§19.13).

19.4 Gateway: gateway-plugin/http_server.py

New module, started/stopped by AndroidAdapter.connect()/disconnect() next to the WS server.

  • Server: http.server.ThreadingHTTPServer + BaseHTTPRequestHandler, run in a daemon thread (one thread per connection — fine at single-user scale). The handler thread never touches adapter state directly; it bridges into the gateway's asyncio loop with asyncio.run_coroutine_threadsafe(coro, loop) (the loop is captured at start, same loop the WS server runs on).
  • Config: ANDROID_HTTP_PORT (default 8791), same bind host as the WS (ANDROID_WS_HOST). Optional TLS via ANDROID_HTTP_CERT/ANDROID_HTTP_KEY (ssl.SSLContext on the server) — same posture as the WS: plaintext on a trusted LAN by default, TLS for remote/Tailscale setups.
  • Bind failure is NON-fatal (unlike the WS): log a warning, disable the HTTP leg, show it in the inspector. The plugin must keep working WS-only.
  • Port-conflict lock: same flock pattern the WS uses (host:port key).

Endpoints

Endpoint Auth Purpose
GET /v1/health none Liveness probe → 200 {"ok": true}. Leaks nothing (no token echo, no device info). The app races this against the WS dial at startup.
POST /v1/frame Bearer token Accept any JSON frame the WS accepts (except binary media). Body = one frame envelope (04-wire-protocol.md). Dispatched through the same adapter handlers as WS (on_message_send, on_search, …).
GET /v1/events?cursor=N Bearer token SSE stream: catch-up from the outbox, then live frames (§19.5).
GET /v1/poll?cursor=N Bearer token Long-poll fallback where SSE is blocked (§19.6).

Auth & limits

  • Authorization: Bearer <token>; verified with the existing constant-time verify_token(); 401 on failure. Device identity via X-Iris-Device header (same device_id the app uses for hello; same allowlist check).
  • Request body cap 64 KiB, Content-Type: application/json enforced (frames are small; media never travels here in v1).
  • Rate limit: token bucket per device, same parameters as the WS inbound limit (INBOUND_RATE_PER_S / INBOUND_BURST); 429 on exceed.
  • No CORS headers (app clients only); unknown paths → 404.

19.5 SSE stream design (GET /v1/events)

Wire format (standard SSE, three fields):

id: 1043
event: frame
data: {"v":1,"type":"message","chat_id":"android:default",...}

: hb                      ← comment heartbeat every 15 s (keeps proxies alive)
  • id = outbox cursor. This is what makes resume trivial: on reconnect the client sends Last-Event-ID (or ?cursor=) and the server replays outbox.replay(cursor) — exactly the sync semantics, no new machinery.
  • Stream open sequence:
    1. Replay outbox rows with cursor > N (bounded by the existing _REPLAY_LIMIT), each as an event: frame with its id.
    2. One event: hello carrying the hello.ack payload (server_caps, sync_cursor, last_pushed_cursor, channels) — the HTTP equivalent of pairing-ack; the app treats it like hello.ack.
    3. Live frames as they are produced.
  • Live fan-out hook: in adapter._broadcast_or_log, after outbox.append() returns the cursor, push (cursor, frame_json) into every live HTTP subscriber's queue. The direct status broadcasts (ws_server.broadcast(protocol.status(...))) get a second fan-out call with cursor = None (SSE event without id).
  • Thread model: each SSE connection owns its handler thread, which blocks on a cross-thread queue.get() (via run_coroutine_threadsafe, 30 s timeout → write : hb and loop) and writes to wfile + flush().
  • Backpressure: bounded queue (256). A subscriber that can't keep up is dropped; the client reconnects with Last-Event-ID and catches up from the outbox. Single-user scale makes this a non-event in practice.
  • App-side reader: hand-rolled over OkHttp's streaming ResponseBody (read lines; id: / event: / data:; blank line = dispatch). ~100 lines, no new dependency. Reconnect with exponential backoff + Last-Event-ID.

19.6 Long-poll fallback (GET /v1/poll)

For networks/proxies that buffer or kill SSE:

  • GET /v1/poll?cursor=N → server holds the request (asyncio waiter on the subscriber queue) until a frame with cursor > N exists or 25 s pass.
  • Response: 200 {"cursor": <new high-water>, "frames": [ ... ]} (frames may be empty on timeout; the app immediately re-polls with the new cursor).
  • The app switches to long-poll automatically after two consecutive SSE open failures, and back to SSE on the next full (re)connect.

19.7 Request/response over HTTP

POST /v1/frame is accept-and-ack:

  • 202 {"ok": true} — frame accepted and dispatched.
  • 4xx with an error frame as the JSON body for validation rejections (empty message, automation-channel read-only, rate limit → 429, bad JSON → 400). These are the same error frames the WS path sends via send_to; over HTTP they double as the HTTP response.
  • Async responses (user echo, search results, channel.list, the agent reply, streaming updates) arrive on the event stream carrying the same id — the app's existing request-id correlation works unchanged.
  • Consequence: send_to(device_id, …) error replies for HTTP-originated requests are instead broadcast (single-user model; the SSE stream delivers them). The dispatch refactor must tag the origin so WS-originated requests keep point-to-point errors.

19.8 Delivery counting & push interaction (critical)

_broadcast_or_log fires push when delivered == 0. With the HTTP leg, a device reading SSE is a live subscriber:

delivered = await self._ws_server.broadcast(frame)
delivered += await self._http_server.fanout(frame, cursor)   # live SSE/poll subs
...
if delivered == 0:  # → outbox + push (unchanged)

If this is forgotten, every message would push and stream to a device that is already receiving it. Related bookkeeping:

  • has_devices() / status must count HTTP subscribers as connected devices (mark the device's transport ws | http in the connection registry).
  • last_pushed_cursor / notification dedupe (08-push.md §8.8) is unchanged — SSE-replayed frames carry the same cursor envelope as sync-replayed ones, so the app's existing dedupe applies.

19.9 App side

New iris/net/HttpGateway.kt (OkHttp) + a transport state machine inside GatewayClient (or a thin Transport wrapper around it):

  • API: health(timeoutMs), postFrame(json): Result, events(cursor, onFrame, onHello): Job (SSE reader), poll(cursor): Result.

  • State machine:

    State Send path Receive path
    WS_CONNECTED WS frame WS
    HTTP_FALLBACK POST /v1/frame SSE (or long-poll)
    CONNECTING / RECONNECTING queued/dropped as today —
  • On WS loss: switch to HTTP_FALLBACK immediately — open the SSE stream (catch-up from the local cursor is free) and route sends to POST. No backoff gate on the send path; the WS redial loop keeps running in the background.

  • At startup (the key UX fix): race the WS dial against GET /v1/health (2 s timeout). WS dial fails + health OK → straight into HTTP_FALLBACK: the user can send in < 1 s after opening the app, instead of waiting out dial timeouts and backoff.

  • On WS reconnect: close the SSE stream, resume WS-only (lowest latency, media available again).

  • Send path: sendMessage() builds the same message.send frame JSON and writes it to WS or POST depending on state. The State.Connected gate in ChatScreen.doSend() becomes state is Connected || state is HttpFallback.

  • Media: disabled in the composer while in HTTP_FALLBACK (v1).

  • UI: status pill shows "connected" (WS) or "connected · http" (fallback) — both green; the fallback is a healthy state, not an error.

19.10 Security

  • Same token, constant-time verify, same bind host, same device allowlist as the WS (09-pairing-security.md threat model unchanged — the HTTP leg adds no new trust boundary, only a second door with the same lock).
  • /v1/health is unauthenticated by design (it answers "is the gateway alive?"); it must not reflect tokens, device ids, or version strings.
  • TLS: optional, same cert pattern as the WS; plaintext is a LAN-only default, identical to today's WS posture.
  • New attack-surface items to keep small: 64 KiB body cap, strict content-type, per-device rate limit, no directory listing, no CORS.

19.11 Failure modes

Failure Behavior
Gateway fully down Both legs dead → app shows offline; sends queue (app-side outbox, follow-up work) or are dropped with a visible "not sent" state. Push is the wake path when the gateway comes back (08-push.md).
WS down, HTTP up Normal HTTP_FALLBACK operation — text chat fully functional, media paused.
SSE blocked by a proxy Two failures → long-poll loop (§19.6).
HTTP port firewalled, WS up WS-only operation (today's behavior); health fails at startup, no fallback attempted.
Both flaky Existing WS backoff + SSE/poll backoff run independently; outbox + cursor keep both paths idempotent.
Slow SSE subscriber Dropped at queue overflow; reconnects with Last-Event-ID, catches up from outbox.

19.12 Testing

  • Python (hermes-agent/tests/gateway/test_android_http.py, run via scripts/run_tests.sh):
    • auth: bad/missing token → 401; allowlist rejection; constant-time verify reused.
    • POST /v1/frame: valid message.send dispatches (agent turn fires); empty text → 400 error frame; automation channel → 409/400; rate limit → 429.
    • SSE: catch-up rows carry correct ids; event: hello present; a live frame appended after connect arrives on the stream; Last-Event-ID resume replays exactly the delta; heartbeat observed within 15 s.
    • long-poll: returns on new frame; empty 200 at timeout with advanced cursor.
    • delivery counting: frame with only an SSE subscriber → delivered ≥ 1 → no push fired (the critical regression test for §19.8).
  • Probe: ws_probe.py gains an --http mode (health, post, SSE read with assertion flags, per gateway-plugin/tests/README.md).
  • Kotlin (:shared commonTest): SSE parser (multi-line data, comments, Last-Event-ID bookkeeping); transport state machine transitions (fake clock: WS-loss → immediate fallback; startup race → fallback in < 1 s).
  • E2E (e2e.py, new scenario): point the app at a dead WS port with the HTTP leg live → send a message → assert user echo + agent reply arrive via SSE; timing assertion: send → user echo < 1 s on LAN. Live-verify on the device via ADB (screenshot of the "connected · http" pill).

19.13 Non-goals (v1) / future

  • Media over HTTP (v2): POST /v1/media (chunked, same sha256 contract as 07-media.md) + GET /v1/media/{id} for pull/playback. Unblocks attachments in fallback mode.
  • App-side send outbox (companion work, separate doc): queue sends locally when both legs are down; drains over whichever leg recovers. This doc removes the 2–20 s wait; the outbox removes the last "gateway was down for 30 s" data-loss case.
  • QUIC / WebTransport if handover flakiness persists after this + the outbox (connection migration would make the fallback rare).
  • Per-device tokens (16-open-questions.md #3) apply to both legs identically when implemented.

19.14 Effort & change list

Slice Files Est.
Gateway leg new gateway-plugin/http_server.py (~450 lines); adapter.py hooks (start/stop, fan-out in _broadcast_or_log + status path, delivery counting, dispatch-origin tag); protocol.py unchanged 2–3 d
App leg new app/shared/.../net/HttpGateway.kt (SSE reader + poll); GatewayClient.kt state machine + startup race; ChatScreen.kt gate + status pill; composer media-disable in fallback 2–3 d
Tests + e2e + docs per §19.12; frames.schema.json unchanged (no new frame types); 09-pairing-security.md + 13-testing.md cross-references 1–2 d

Total: ~1 week, each slice independently shippable (gateway leg is inert until the app uses it; app leg degrades to today's behavior if the HTTP port is closed).