# 19 — HTTP Fallback Transport (the "HTTP leg") A second, **short-lived-connection** transport next to the WebSocket: the same JSON frames, the same outbox/cursor, the same token — served over plain HTTP by the gateway. When the WS is down (flaky network, NAT timeout, app just relaunched), the app **sends over `POST` and receives over SSE** instead of waiting 2–20 s for a WS redial. Status: **implemented — and now the ONLY transport.** The WebSocket leg has been removed entirely from the codebase (gateway `ws_server.py` deleted; WS-only frames `hello`/`ping`/`pong`/`media.upload.*`/`media.pull*` dropped from `protocol.py` and `Protocol.kt`; `GatewayClient` is HTTP-only with no `HttpFallback` state — it *is* the connected state). HTTP is the primary and sole transport: v1 (JSON frames over POST/SSE/long-poll) and v2 (media over `POST /v1/media` + `GET /v1/media/{id}`, §19.15). Gateway leg: `gateway-plugin/http_server.py` (+ `dispatch.py` for frame dispatch); app leg: `app/shared/src/commonMain/kotlin/iris/net/HttpGateway.kt` + `GatewayClient.kt`. Legacy `ws(s)://` URLs entered by users are still accepted and rewritten to `http(s)://` (`HttpGateway.deriveHttpUrl`). Complements — does not replace — `04-wire-protocol.md` (frames), `08-push.md` (outbox/sync/push), and `09-pairing-security.md` (auth model). > **Note:** the rest of this document describes the original design, in > which HTTP was a *fallback* next to a WS primary. That framing is > historical; where it says "WS (primary)" / "HTTP (fallback)", read > "HTTP (the only transport)". ## 19.1 Problem Today the WS is the *only* transport, and the app hard-gates sending on a live socket (`ChatScreen.doSend()` no-ops unless `State.Connected`; `GatewayClient.sendMessage()` drops when `socket == null`). Consequences: - **App killed → reopened:** full cold dial (TCP + TLS + `hello`/`hello.ack`, 15 s dial timeout) before the user can send. On a flaky network the first dial often fails → backoff → second dial. Observed: 2–20 s of "can't send". - **Long-lived WS is the most fragile connection type on mobile:** idle sockets expire in router/CGNAT NAT tables, die on WiFi↔cellular handover, and are killed aggressively by OEM power management (MIUI on the test device). There is no foreground service holding the WS. - **Stale detection is slow:** 20 s ping interval, 60 s reap — a dead-but- unclosed socket can sit for up to a minute before redial. ## 19.2 Why HTTP (and why not the alternatives) Short-lived HTTP requests are dramatically more resilient on mobile networks than a long-lived socket: no NAT table entry to expire, no proxy idle-kill, each request is a fresh connection (fast with TLS resumption), and they work through the restrictive proxies that mangle WebSockets. Sending a message becomes a single `POST` that completes in well under a second on a LAN — independent of whether the WS is up. Alternatives considered and rejected (research, 2026-08): | Option | Verdict | | --- | --- | | **MQTT broker** (QoS 1, persistent sessions) | Best protocol for flaky links, but new infra (broker process) + new Python dep (`paho-mqtt`, breaks the zero-new-deps rule) + new Kotlin dep + frame↔topic bridge. Overkill for a 1-user agent. | | **ntfy as the send path** (app publishes to a topic the gateway subscribes to) | Adds a third party to the critical send path; public ntfy.sh is already known-flaky. Not worth it. | | **WebTransport / QUIC** | The *real* fix for handover flakiness (connection migration), but no OkHttp support and `aioquic` is a new Python dep. Future option if this doc's approach is still not enough. | | **gRPC** | New deps both sides; no advantage over WS+SSE here. | | **Inverted connection** (app runs a local HTTP server, gateway pushes to the phone) | LAN-only, breaks on cellular/remote, security mess. Rejected. | | **Matrix / full chat server** | Massive overkill for a personal agent. | **Zero new Python dependencies is preserved:** the HTTP leg is stdlib `http.server` (a `ThreadingHTTPServer` in a daemon thread) bridged into the gateway's asyncio loop. The app side uses the OkHttp it already depends on (hand-rolled SSE reader — the format is trivial; `okhttp-eventsource` is an acceptable alternative if preferred). ## 19.3 Shape ``` ┌──────────────────────── hermes gateway process ───────────────────────┐ │ AndroidAdapter │ │ │ frames (same protocol.Frame objects) │ │ ▼ │ │ _broadcast_or_log ──► outbox.append(cursor) ──► push (if no live) │ │ │ │ │ │ ▼ ▼ │ │ WsServer (asyncio, :8790) HttpServer (stdlib thread, :8791) │ │ primary: full protocol fallback: POST /v1/frame, │ │ incl. binary media GET /v1/events (SSE), /v1/poll │ └───────────────┬──────────────────────────────┬────────────────────────┘ │ WS (primary) │ HTTP (fallback) ┌─────────┴──────────────────────────────┴────────┐ │ APP: transport state machine │ │ WS up → WS only (media works, lowest latency)│ │ WS down → send via POST, receive via SSE/poll │ └─────────────────────────────────────────────────┘ ``` **v1 scope** | Over HTTP | WS-only | | --- | --- | | All JSON request frames (`message.send`, `search`, `channel.*`, `commands.catalog`, `agent.stop`/`agent.steer`, …) via one generic endpoint | — | | All event/response frames via SSE (or long-poll) | — | | `sync` catch-up (same outbox, same cursor) | — | | Media upload + pull (`POST /v1/media`, `GET /v1/media/{id}`, v2 — §19.15) | — | With v2 the HTTP leg is feature-complete: media no longer needs the WS (the composer's attach button is enabled in `HTTP_FALLBACK` too). The WS binary media frames remain accepted for WS clients, but the app routes media over HTTP whenever the WS is down. ## 19.4 Gateway: `gateway-plugin/http_server.py` New module, started/stopped by `AndroidAdapter.connect()`/`disconnect()` next to the WS server. - **Server:** `http.server.ThreadingHTTPServer` + `BaseHTTPRequestHandler`, run in a **daemon thread** (one thread per connection — fine at single-user scale). The handler thread never touches adapter state directly; it bridges into the gateway's asyncio loop with `asyncio.run_coroutine_threadsafe(coro, loop)` (the loop is captured at start, same loop the WS server runs on). - **Config:** `ANDROID_HTTP_PORT` (default **8791**), same bind host as the WS (`ANDROID_WS_HOST`). Optional TLS via `ANDROID_HTTP_CERT`/`ANDROID_HTTP_KEY` (`ssl.SSLContext` on the server) — same posture as the WS: plaintext on a trusted LAN by default, TLS for remote/Tailscale setups. - **Bind failure is NON-fatal** (unlike the WS): log a warning, disable the HTTP leg, show it in the inspector. The plugin must keep working WS-only. - Port-conflict lock: same flock pattern the WS uses (`host:port` key). ### Endpoints | Endpoint | Auth | Purpose | | --- | --- | --- | | `GET /v1/health` | none | Liveness probe → `200 {"ok": true}`. Leaks nothing (no token echo, no device info). The app races this against the WS dial at startup. | | `POST /v1/frame` | Bearer token | Accept **any** JSON frame the WS accepts (except binary media). Body = one frame envelope (`04-wire-protocol.md`). Dispatched through the *same* adapter handlers as WS (`on_message_send`, `on_search`, …). | | `GET /v1/events?cursor=N` | Bearer token | **SSE** stream: catch-up from the outbox, then live frames (§19.5). | | `GET /v1/poll?cursor=N` | Bearer token | **Long-poll** fallback where SSE is blocked (§19.6). | | `POST /v1/media` | Bearer token | **Media upload** (v2, §19.15): whole file as the body, metadata in `X-Iris-Media-*` headers. | | `GET /v1/media/{media_id}` | Bearer token | **Media pull** (v2, §19.15): streams an outbound offer as the response body. | ### Auth & limits - `Authorization: Bearer `; verified with the existing constant-time `verify_token()`; `401` on failure. Device identity via `X-Iris-Device` header (same `device_id` the app uses for `hello`; same allowlist check). - Request body cap **64 KiB**, `Content-Type: application/json` enforced (frames are small; media never travels here in v1). - Rate limit: token bucket per device, same parameters as the WS inbound limit (`INBOUND_RATE_PER_S` / `INBOUND_BURST`); `429` on exceed. - No CORS headers (app clients only); unknown paths → `404`. ## 19.5 SSE stream design (`GET /v1/events`) Wire format (standard SSE, three fields): ``` id: 1043 event: frame data: {"v":1,"type":"message","chat_id":"android:default",...} : hb ← comment heartbeat every 15 s (keeps proxies alive) ``` - **`id` = outbox cursor.** This is what makes resume trivial: on reconnect the client sends `Last-Event-ID` (or `?cursor=`) and the server replays `outbox.replay(cursor)` — exactly the `sync` semantics, no new machinery. - **Stream open sequence:** 1. Replay outbox rows with `cursor > N` (bounded by the existing `_REPLAY_LIMIT`), each as an `event: frame` with its `id`. 2. One `event: hello` carrying the `hello.ack` payload (`server_caps`, `sync_cursor`, `last_pushed_cursor`, `channels`) — the HTTP equivalent of pairing-ack; the app treats it like `hello.ack`. 3. Live frames as they are produced. - **Live fan-out hook:** in `adapter._broadcast_or_log`, after `outbox.append()` returns the cursor, push `(cursor, frame_json)` into every live HTTP subscriber's queue. The direct `status` broadcasts (`ws_server.broadcast(protocol.status(...))`) get a second fan-out call with `cursor = None` (SSE event without `id`). - **Thread model:** each SSE connection owns its handler thread, which blocks on a cross-thread `queue.get()` (via `run_coroutine_threadsafe`, 30 s timeout → write `: hb` and loop) and writes to `wfile` + `flush()`. - **Backpressure:** bounded queue (256). A subscriber that can't keep up is dropped; the client reconnects with `Last-Event-ID` and catches up from the outbox. Single-user scale makes this a non-event in practice. - **App-side reader:** hand-rolled over OkHttp's streaming `ResponseBody` (read lines; `id:` / `event:` / `data:`; blank line = dispatch). ~100 lines, no new dependency. Reconnect with exponential backoff + `Last-Event-ID`. ## 19.6 Long-poll fallback (`GET /v1/poll`) For networks/proxies that buffer or kill SSE: - `GET /v1/poll?cursor=N` → server holds the request (asyncio waiter on the subscriber queue) until a frame with `cursor > N` exists or **25 s** pass. - Response: `200 {"cursor": , "frames": [ ... ]}` (frames may be empty on timeout; the app immediately re-polls with the new cursor). - The app switches to long-poll automatically after **two consecutive SSE open failures**, and back to SSE on the next full (re)connect. ## 19.7 Request/response over HTTP `POST /v1/frame` is accept-and-ack: - `202 {"ok": true}` — frame accepted and dispatched. - `4xx` with an **error frame as the JSON body** for validation rejections (empty message, automation-channel read-only, rate limit → `429`, bad JSON → `400`). These are the same `error` frames the WS path sends via `send_to`; over HTTP they double as the HTTP response. - **Async responses** (user echo, `search` results, `channel.list`, the agent reply, streaming updates) arrive on the **event stream** carrying the same `id` — the app's existing request-id correlation works unchanged. - Consequence: `send_to(device_id, …)` error replies for HTTP-originated requests are instead **broadcast** (single-user model; the SSE stream delivers them). The dispatch refactor must tag the origin so WS-originated requests keep point-to-point errors. ## 19.8 Delivery counting & push interaction (critical) `_broadcast_or_log` fires push when `delivered == 0`. With the HTTP leg, a device reading SSE **is** a live subscriber: ```python delivered = await self._ws_server.broadcast(frame) delivered += await self._http_server.fanout(frame, cursor) # live SSE/poll subs ... if delivered == 0: # → outbox + push (unchanged) ``` If this is forgotten, every message would push *and* stream to a device that is already receiving it. Related bookkeeping: - `has_devices()` / `status` must count HTTP subscribers as connected devices (mark the device's transport `ws` | `http` in the connection registry). - `last_pushed_cursor` / notification dedupe (`08-push.md` §8.8) is unchanged — SSE-replayed frames carry the same `cursor` envelope as `sync`-replayed ones, so the app's existing dedupe applies. ## 19.9 App side New `iris/net/HttpGateway.kt` (OkHttp) + a transport state machine inside `GatewayClient` (or a thin `Transport` wrapper around it): - **API:** `health(timeoutMs)`, `postFrame(json): Result`, `events(cursor, onFrame, onHello): Job` (SSE reader), `poll(cursor): Result`. - **State machine:** | State | Send path | Receive path | | --- | --- | --- | | `WS_CONNECTED` | WS frame | WS | | `HTTP_FALLBACK` | `POST /v1/frame` | SSE (or long-poll) | | `CONNECTING` / `RECONNECTING` | queued/dropped as today | — | - **On WS loss:** switch to `HTTP_FALLBACK` **immediately** — open the SSE stream (catch-up from the local cursor is free) and route sends to POST. No backoff gate on the send path; the WS redial loop keeps running in the background. - **At startup (the key UX fix):** race the WS dial against `GET /v1/health` (2 s timeout). WS dial fails + health OK → straight into `HTTP_FALLBACK`: the user can send in **< 1 s** after opening the app, instead of waiting out dial timeouts and backoff. - **On WS reconnect:** close the SSE stream, resume WS-only (lowest latency, media available again). - **Send path:** `sendMessage()` builds the same `message.send` frame JSON and writes it to WS or POST depending on state. The `State.Connected` gate in `ChatScreen.doSend()` becomes `state is Connected || state is HttpFallback`. - **Media:** works in `HTTP_FALLBACK` too (v2, §19.15) — uploads go via `POST /v1/media`, pulls via `GET /v1/media/{id}`; the composer's attach button is enabled in both connected states. - **UI:** status pill shows "connected" (WS) or "connected · http" (fallback) — both green; the fallback is a healthy state, not an error. ## 19.10 Security - Same token, constant-time verify, same bind host, same device allowlist as the WS (`09-pairing-security.md` threat model unchanged — the HTTP leg adds no new trust boundary, only a second door with the same lock). - `/v1/health` is unauthenticated by design (it answers "is the gateway alive?"); it must not reflect tokens, device ids, or version strings. - TLS: optional, same cert pattern as the WS; plaintext is a LAN-only default, identical to today's WS posture. - New attack-surface items to keep small: 64 KiB body cap, strict content-type, per-device rate limit, no directory listing, no CORS. ## 19.11 Failure modes | Failure | Behavior | | --- | --- | | Gateway fully down | Both legs dead → app shows offline; sends queue (app-side outbox, follow-up work) or are dropped with a visible "not sent" state. Push is the wake path when the gateway comes back (`08-push.md`). | | WS down, HTTP up | Normal `HTTP_FALLBACK` operation — text chat fully functional, media paused. | | SSE blocked by a proxy | Two failures → long-poll loop (§19.6). | | HTTP port firewalled, WS up | WS-only operation (today's behavior); `health` fails at startup, no fallback attempted. | | Both flaky | Existing WS backoff + SSE/poll backoff run independently; outbox + cursor keep both paths idempotent. | | Slow SSE subscriber | Dropped at queue overflow; reconnects with `Last-Event-ID`, catches up from outbox. | ## 19.12 Testing - **Python** (`hermes-agent/tests/gateway/test_android_http.py`, run via `scripts/run_tests.sh`): - auth: bad/missing token → 401; allowlist rejection; constant-time verify reused. - `POST /v1/frame`: valid `message.send` dispatches (agent turn fires); empty text → 400 error frame; automation channel → 409/400; rate limit → 429. - SSE: catch-up rows carry correct `id`s; `event: hello` present; a live frame appended after connect arrives on the stream; `Last-Event-ID` resume replays exactly the delta; heartbeat observed within 15 s. - long-poll: returns on new frame; empty 200 at timeout with advanced cursor. - **delivery counting:** frame with only an SSE subscriber → `delivered ≥ 1` → **no push fired** (the critical regression test for §19.8). - **media (v2, §19.15):** `POST /v1/media` happy path (201 ack + cached entry), sha256 mismatch, oversize → 413, missing ref / bad kind → 400, auth → 401, magic-byte reclassification; `GET /v1/media/{id}` happy path (bytes + content-type), unknown id → 404, denied path → 404. - **Probe:** `ws_probe.py` gains an `--http` mode (health, post, SSE read with assertion flags, per `gateway-plugin/tests/README.md`) + `--http-media FILE` (v2: upload round-trip via `POST /v1/media`, exit 23 on rejection). - **Kotlin** (`:shared` commonTest): SSE parser (multi-line data, comments, `Last-Event-ID` bookkeeping); transport state machine transitions (fake clock: WS-loss → immediate fallback; startup race → fallback in < 1 s). - **E2E** (`e2e.py`, new scenario): point the app at a dead WS port with the HTTP leg live → send a message → assert user echo + agent reply arrive via SSE; timing assertion: send → user echo < 1 s on LAN. Live-verify on the device via ADB (screenshot of the "connected · http" pill). ## 19.13 Non-goals (v1) / future - ~~**Media over HTTP** (v2)~~ — **done** (§19.15): `POST /v1/media` (whole-file body, sha256 contract per `07-media.md`) + `GET /v1/media/{id}` for pull/playback. Attachments work in fallback mode. - **App-side send outbox** (companion work, separate doc): queue sends locally when *both* legs are down; drains over whichever leg recovers. This doc removes the 2–20 s wait; the outbox removes the last "gateway was down for 30 s" data-loss case. - **QUIC / WebTransport** if handover flakiness persists after this + the outbox (connection migration would make the fallback rare). - Per-device tokens (`16-open-questions.md` #3) apply to both legs identically when implemented. ## 19.15 Media over HTTP (v2) The last WS-only feature, closed out so the HTTP leg is feature-complete. Same contracts as `07-media.md` — only the transport changes. ### Upload — `POST /v1/media` One request per file (no chunked/resumable protocol — HTTP handles the body; single-user scale makes resume unnecessary): ``` POST /v1/media Authorization: Bearer X-Iris-Device: X-Iris-Media-Ref: up_123456 # app-chosen ref (mu_*/up_*), ≤ 64 chars X-Iris-Media-Kind: image|audio|video|document|voice X-Iris-Media-Filename: photo.jpg X-Iris-Media-Sha256: <64 hex> # precomputed (headers precede the body) Content-Type: # doubles as the declared MIME Content-Length: ``` - **Response:** `201` with the `media.upload.ack` frame as the body (`{ok, media_ref}`); validation failures return the `error` frame as the 4xx body with the same codes as the WS path (`media_too_large` → 413, `unsupported` → 400, `internal` → 500, `not_found` → 404). - **Server flow:** the body is streamed to a temp file in 256 KiB reads (bounded RAM, same `UploadSession` as the WS path), then `complete_upload` verifies size + sha256, re-sniffs the kind from magic bytes (the client's declared kind is not trusted), and caches via the hermes `cache_*_from_bytes` helpers. Runs entirely on the handler thread — no asyncio bridge (plain file IO). - **Limits:** `Content-Length` is checked against `max_upload_bytes` *before* reading the body (early 413); the 64 KiB `/v1/frame` body cap does not apply. Same per-device rate limit as the other endpoints. - **Abort:** a client that disconnects mid-body leaves a short read → the upload session (temp file) is discarded; nothing is cached. - The ref then travels in `message.send`'s `media_refs` exactly as on the WS path (single-use, resolved to `MessageEvent.media_urls`). ### Pull — `GET /v1/media/{media_id}` - `media.offer` is a plain JSON event frame — it arrives on the SSE stream unchanged; only the byte transfer moves to HTTP. - **Response:** `200` with the file as the body, `Content-Type: `, `Content-Length: `, `Content-Disposition: attachment; filename=""`. Unknown id or a path that fails delivery validation → `404` with the `error` frame (`not_found`) — the delivery-path check is re-run at pull time, exactly as the WS `media.pull` handler does. - The app streams the body into its media cache (same `MediaCache.openWriter` path as the WS pull); playback is unchanged (`07-media.md` §7.4). ### What stays WS-only Nothing feature-wise. The WS binary media frames (`media.upload.start/end`, `media.pull` + binary chunks) remain accepted for WS clients, and `POST /v1/frame` still rejects those frame types (they have HTTP endpoint equivalents now, not a WS dependency). ## 19.14 Effort & change list | Slice | Files | Est. | | --- | --- | --- | | Gateway leg | new `gateway-plugin/http_server.py` (~450 lines); `adapter.py` hooks (start/stop, fan-out in `_broadcast_or_log` + status path, delivery counting, dispatch-origin tag); `protocol.py` unchanged | 2–3 d | | App leg | new `app/shared/.../net/HttpGateway.kt` (SSE reader + poll); `GatewayClient.kt` state machine + startup race; `ChatScreen.kt` gate + status pill; composer media-disable in fallback | 2–3 d | | Tests + e2e + docs | per §19.12; `frames.schema.json` unchanged (no new frame types); `09-pairing-security.md` + `13-testing.md` cross-references | 1–2 d | Total: **~1 week**, each slice independently shippable (gateway leg is inert until the app uses it; app leg degrades to today's behavior if the HTTP port is closed).