HTTP fallback leg (docs/19): POST /v1/frame + SSE /v1/events + long-poll

Second, short-lived-connection transport next to the WS: same frames,
same outbox/cursor, same token, served over plain HTTP (stdlib
ThreadingHTTPServer bridged into the asyncio loop; zero new deps).

- http_server.py: /v1/health (unauthenticated), POST /v1/frame
  (accept-and-ack; validation rejections as 4xx error frames), SSE
  /v1/events (outbox catch-up with id=cursor, event: hello, 15s
  heartbeat, bounded-queue backpressure), long-poll /v1/poll (25s hold).
  Bearer token + X-Iris-Device (same allowlist as WS hello), 64 KiB body
  cap, per-device rate limit, optional TLS, non-fatal bind failure.
- ws_server.py: inbound dispatch chain extracted to shared
  dispatch_frame() used by both transports.
- adapter.py: ANDROID_HTTP_PORT/CERT/KEY config; start/stop next to the
  WS; delivery counting in _broadcast_or_log (an SSE subscriber is a
  live subscriber -> no push, docs/19 19.8); _reply() routes
  point-to-point replies into the in-flight HTTP response (reply sink)
  or broadcasts when the device has no live WS (19.7); status/typing/
  channel events fan out to both transports.
- ws_probe.py: --http mode (health + POST + SSE turn drive, same
  assertion flags); tests/README updated.
- Tests: hermes-agent/tests/gateway/test_android_http.py (23 tests,
  incl. the 19.8 delivery-counting regression); test_android.py (74)
  still green.
This commit is contained in:
ARIA committed 2026-08-22 14:10:14 +02:00
1 parent 524ed8ce53
commit e5c7d690b8
7 files changed
+1315 -89

No files matched your search

+318
View File
@@ -0,0 +1,318 @@
# 19 — HTTP Fallback Transport (the "HTTP leg")
A second, **short-lived-connection** transport next to the WebSocket: the same
JSON frames, the same outbox/cursor, the same token — served over plain HTTP
by the gateway. When the WS is down (flaky network, NAT timeout, app just
relaunched), the app **sends over `POST` and receives over SSE** instead of
waiting 2–20 s for a WS redial.
Status: **design (proposed, not built)**. Complements — does not replace —
`04-wire-protocol.md` (frames), `08-push.md` (outbox/sync/push), and
`09-pairing-security.md` (auth model).
## 19.1 Problem
Today the WS is the *only* transport, and the app hard-gates sending on a live
socket (`ChatScreen.doSend()` no-ops unless `State.Connected`;
`GatewayClient.sendMessage()` drops when `socket == null`). Consequences:
- **App killed → reopened:** full cold dial (TCP + TLS + `hello`/`hello.ack`,
15 s dial timeout) before the user can send. On a flaky network the first
dial often fails → backoff → second dial. Observed: 2–20 s of "can't send".
- **Long-lived WS is the most fragile connection type on mobile:** idle
sockets expire in router/CGNAT NAT tables, die on WiFi↔cellular handover,
and are killed aggressively by OEM power management (MIUI on the test
device). There is no foreground service holding the WS.
- **Stale detection is slow:** 20 s ping interval, 60 s reap — a dead-but-
unclosed socket can sit for up to a minute before redial.
## 19.2 Why HTTP (and why not the alternatives)
Short-lived HTTP requests are dramatically more resilient on mobile networks
than a long-lived socket: no NAT table entry to expire, no proxy idle-kill,
each request is a fresh connection (fast with TLS resumption), and they work
through the restrictive proxies that mangle WebSockets. Sending a message
becomes a single `POST` that completes in well under a second on a LAN —
independent of whether the WS is up.
Alternatives considered and rejected (research, 2026-08):
| Option | Verdict |
| --- | --- |
| **MQTT broker** (QoS 1, persistent sessions) | Best protocol for flaky links, but new infra (broker process) + new Python dep (`paho-mqtt`, breaks the zero-new-deps rule) + new Kotlin dep + frame↔topic bridge. Overkill for a 1-user agent. |
| **ntfy as the send path** (app publishes to a topic the gateway subscribes to) | Adds a third party to the critical send path; public ntfy.sh is already known-flaky. Not worth it. |
| **WebTransport / QUIC** | The *real* fix for handover flakiness (connection migration), but no OkHttp support and `aioquic` is a new Python dep. Future option if this doc's approach is still not enough. |
| **gRPC** | New deps both sides; no advantage over WS+SSE here. |
| **Inverted connection** (app runs a local HTTP server, gateway pushes to the phone) | LAN-only, breaks on cellular/remote, security mess. Rejected. |
| **Matrix / full chat server** | Massive overkill for a personal agent. |
**Zero new Python dependencies is preserved:** the HTTP leg is stdlib
`http.server` (a `ThreadingHTTPServer` in a daemon thread) bridged into the
gateway's asyncio loop. The app side uses the OkHttp it already depends on
(hand-rolled SSE reader — the format is trivial; `okhttp-eventsource` is an
acceptable alternative if preferred).
## 19.3 Shape
```
┌──────────────────────── hermes gateway process ───────────────────────┐
│ AndroidAdapter │
│ │ frames (same protocol.Frame objects) │
│ ▼ │
│ _broadcast_or_log ──► outbox.append(cursor) ──► push (if no live) │
│ │ │ │
│ ▼ ▼ │
│ WsServer (asyncio, :8790) HttpServer (stdlib thread, :8791) │
│ primary: full protocol fallback: POST /v1/frame, │
│ incl. binary media GET /v1/events (SSE), /v1/poll │
└───────────────┬──────────────────────────────┬────────────────────────┘
│ WS (primary) │ HTTP (fallback)
┌─────────┴──────────────────────────────┴────────┐
│ APP: transport state machine │
│ WS up → WS only (media works, lowest latency)│
│ WS down → send via POST, receive via SSE/poll │
└─────────────────────────────────────────────────┘
```
**v1 scope**
| Over HTTP (v1) | WS-only (v1) |
| --- | --- |
| All JSON request frames (`message.send`, `search`, `channel.*`, `commands.catalog`, `agent.stop`/`agent.steer`, …) via one generic endpoint | Binary media upload (chunked binary frames) |
| All event/response frames via SSE (or long-poll) | Binary media pull stream |
| `sync` catch-up (same outbox, same cursor) | — |
Media stays WS-only in v1: it is the one part of the protocol that is
inherently binary/streaming, and attachments are a rarer action than sending
text. While in HTTP-fallback mode the composer disables the attach button
("media needs the live connection"). HTTP media endpoints are a v2 item
(§19.13).
## 19.4 Gateway: `gateway-plugin/http_server.py`
New module, started/stopped by `AndroidAdapter.connect()`/`disconnect()` next
to the WS server.
- **Server:** `http.server.ThreadingHTTPServer` + `BaseHTTPRequestHandler`,
run in a **daemon thread** (one thread per connection — fine at single-user
scale). The handler thread never touches adapter state directly; it bridges
into the gateway's asyncio loop with
`asyncio.run_coroutine_threadsafe(coro, loop)` (the loop is captured at
start, same loop the WS server runs on).
- **Config:** `ANDROID_HTTP_PORT` (default **8791**), same bind host as the WS
(`ANDROID_WS_HOST`). Optional TLS via `ANDROID_HTTP_CERT`/`ANDROID_HTTP_KEY`
(`ssl.SSLContext` on the server) — same posture as the WS: plaintext on a
trusted LAN by default, TLS for remote/Tailscale setups.
- **Bind failure is NON-fatal** (unlike the WS): log a warning, disable the
HTTP leg, show it in the inspector. The plugin must keep working WS-only.
- Port-conflict lock: same flock pattern the WS uses (`host:port` key).
### Endpoints
| Endpoint | Auth | Purpose |
| --- | --- | --- |
| `GET /v1/health` | none | Liveness probe → `200 {"ok": true}`. Leaks nothing (no token echo, no device info). The app races this against the WS dial at startup. |
| `POST /v1/frame` | Bearer token | Accept **any** JSON frame the WS accepts (except binary media). Body = one frame envelope (`04-wire-protocol.md`). Dispatched through the *same* adapter handlers as WS (`on_message_send`, `on_search`, …). |
| `GET /v1/events?cursor=N` | Bearer token | **SSE** stream: catch-up from the outbox, then live frames (§19.5). |
| `GET /v1/poll?cursor=N` | Bearer token | **Long-poll** fallback where SSE is blocked (§19.6). |
### Auth & limits
- `Authorization: Bearer <token>`; verified with the existing constant-time
`verify_token()`; `401` on failure. Device identity via `X-Iris-Device`
header (same `device_id` the app uses for `hello`; same allowlist check).
- Request body cap **64 KiB**, `Content-Type: application/json` enforced
(frames are small; media never travels here in v1).
- Rate limit: token bucket per device, same parameters as the WS inbound
limit (`INBOUND_RATE_PER_S` / `INBOUND_BURST`); `429` on exceed.
- No CORS headers (app clients only); unknown paths → `404`.
## 19.5 SSE stream design (`GET /v1/events`)
Wire format (standard SSE, three fields):
```
id: 1043
event: frame
data: {"v":1,"type":"message","chat_id":"android:default",...}
: hb ← comment heartbeat every 15 s (keeps proxies alive)
```
- **`id` = outbox cursor.** This is what makes resume trivial: on reconnect
the client sends `Last-Event-ID` (or `?cursor=`) and the server replays
`outbox.replay(cursor)` — exactly the `sync` semantics, no new machinery.
- **Stream open sequence:**
1. Replay outbox rows with `cursor > N` (bounded by the existing
`_REPLAY_LIMIT`), each as an `event: frame` with its `id`.
2. One `event: hello` carrying the `hello.ack` payload
(`server_caps`, `sync_cursor`, `last_pushed_cursor`, `channels`) — the
HTTP equivalent of pairing-ack; the app treats it like `hello.ack`.
3. Live frames as they are produced.
- **Live fan-out hook:** in `adapter._broadcast_or_log`, after
`outbox.append()` returns the cursor, push `(cursor, frame_json)` into every
live HTTP subscriber's queue. The direct `status` broadcasts
(`ws_server.broadcast(protocol.status(...))`) get a second fan-out call
with `cursor = None` (SSE event without `id`).
- **Thread model:** each SSE connection owns its handler thread, which blocks
on a cross-thread `queue.get()` (via `run_coroutine_threadsafe`, 30 s
timeout → write `: hb` and loop) and writes to `wfile` + `flush()`.
- **Backpressure:** bounded queue (256). A subscriber that can't keep up is
dropped; the client reconnects with `Last-Event-ID` and catches up from the
outbox. Single-user scale makes this a non-event in practice.
- **App-side reader:** hand-rolled over OkHttp's streaming `ResponseBody`
(read lines; `id:` / `event:` / `data:`; blank line = dispatch). ~100 lines,
no new dependency. Reconnect with exponential backoff + `Last-Event-ID`.
## 19.6 Long-poll fallback (`GET /v1/poll`)
For networks/proxies that buffer or kill SSE:
- `GET /v1/poll?cursor=N` → server holds the request (asyncio waiter on the
subscriber queue) until a frame with `cursor > N` exists or **25 s** pass.
- Response: `200 {"cursor": <new high-water>, "frames": [ ... ]}` (frames may
be empty on timeout; the app immediately re-polls with the new cursor).
- The app switches to long-poll automatically after **two consecutive SSE
open failures**, and back to SSE on the next full (re)connect.
## 19.7 Request/response over HTTP
`POST /v1/frame` is accept-and-ack:
- `202 {"ok": true}` — frame accepted and dispatched.
- `4xx` with an **error frame as the JSON body** for validation rejections
(empty message, automation-channel read-only, rate limit → `429`, bad JSON
→ `400`). These are the same `error` frames the WS path sends via
`send_to`; over HTTP they double as the HTTP response.
- **Async responses** (user echo, `search` results, `channel.list`, the agent
reply, streaming updates) arrive on the **event stream** carrying the same
`id` — the app's existing request-id correlation works unchanged.
- Consequence: `send_to(device_id, …)` error replies for HTTP-originated
requests are instead **broadcast** (single-user model; the SSE stream
delivers them). The dispatch refactor must tag the origin so WS-originated
requests keep point-to-point errors.
## 19.8 Delivery counting & push interaction (critical)
`_broadcast_or_log` fires push when `delivered == 0`. With the HTTP leg, a
device reading SSE **is** a live subscriber:
```python
delivered = await self._ws_server.broadcast(frame)
delivered += await self._http_server.fanout(frame, cursor) # live SSE/poll subs
...
if delivered == 0: # → outbox + push (unchanged)
```
If this is forgotten, every message would push *and* stream to a device that
is already receiving it. Related bookkeeping:
- `has_devices()` / `status` must count HTTP subscribers as connected devices
(mark the device's transport `ws` | `http` in the connection registry).
- `last_pushed_cursor` / notification dedupe (`08-push.md` §8.8) is
unchanged — SSE-replayed frames carry the same `cursor` envelope as
`sync`-replayed ones, so the app's existing dedupe applies.
## 19.9 App side
New `iris/net/HttpGateway.kt` (OkHttp) + a transport state machine inside
`GatewayClient` (or a thin `Transport` wrapper around it):
- **API:** `health(timeoutMs)`, `postFrame(json): Result`,
`events(cursor, onFrame, onHello): Job` (SSE reader), `poll(cursor): Result`.
- **State machine:**
| State | Send path | Receive path |
| --- | --- | --- |
| `WS_CONNECTED` | WS frame | WS |
| `HTTP_FALLBACK` | `POST /v1/frame` | SSE (or long-poll) |
| `CONNECTING` / `RECONNECTING` | queued/dropped as today | — |
- **On WS loss:** switch to `HTTP_FALLBACK` **immediately** — open the SSE
stream (catch-up from the local cursor is free) and route sends to POST.
No backoff gate on the send path; the WS redial loop keeps running in the
background.
- **At startup (the key UX fix):** race the WS dial against
`GET /v1/health` (2 s timeout). WS dial fails + health OK → straight into
`HTTP_FALLBACK`: the user can send in **< 1 s** after opening the app,
instead of waiting out dial timeouts and backoff.
- **On WS reconnect:** close the SSE stream, resume WS-only (lowest latency,
media available again).
- **Send path:** `sendMessage()` builds the same `message.send` frame JSON and
writes it to WS or POST depending on state. The `State.Connected` gate in
`ChatScreen.doSend()` becomes `state is Connected || state is HttpFallback`.
- **Media:** disabled in the composer while in `HTTP_FALLBACK` (v1).
- **UI:** status pill shows "connected" (WS) or "connected · http" (fallback)
— both green; the fallback is a healthy state, not an error.
## 19.10 Security
- Same token, constant-time verify, same bind host, same device allowlist as
the WS (`09-pairing-security.md` threat model unchanged — the HTTP leg adds
no new trust boundary, only a second door with the same lock).
- `/v1/health` is unauthenticated by design (it answers "is the gateway
alive?"); it must not reflect tokens, device ids, or version strings.
- TLS: optional, same cert pattern as the WS; plaintext is a LAN-only
default, identical to today's WS posture.
- New attack-surface items to keep small: 64 KiB body cap, strict
content-type, per-device rate limit, no directory listing, no CORS.
## 19.11 Failure modes
| Failure | Behavior |
| --- | --- |
| Gateway fully down | Both legs dead → app shows offline; sends queue (app-side outbox, follow-up work) or are dropped with a visible "not sent" state. Push is the wake path when the gateway comes back (`08-push.md`). |
| WS down, HTTP up | Normal `HTTP_FALLBACK` operation — text chat fully functional, media paused. |
| SSE blocked by a proxy | Two failures → long-poll loop (§19.6). |
| HTTP port firewalled, WS up | WS-only operation (today's behavior); `health` fails at startup, no fallback attempted. |
| Both flaky | Existing WS backoff + SSE/poll backoff run independently; outbox + cursor keep both paths idempotent. |
| Slow SSE subscriber | Dropped at queue overflow; reconnects with `Last-Event-ID`, catches up from outbox. |
## 19.12 Testing
- **Python** (`hermes-agent/tests/gateway/test_android_http.py`, run via
`scripts/run_tests.sh`):
- auth: bad/missing token → 401; allowlist rejection; constant-time verify reused.
- `POST /v1/frame`: valid `message.send` dispatches (agent turn fires);
empty text → 400 error frame; automation channel → 409/400; rate limit → 429.
- SSE: catch-up rows carry correct `id`s; `event: hello` present; a live
frame appended after connect arrives on the stream; `Last-Event-ID`
resume replays exactly the delta; heartbeat observed within 15 s.
- long-poll: returns on new frame; empty 200 at timeout with advanced cursor.
- **delivery counting:** frame with only an SSE subscriber → `delivered ≥ 1`
→ **no push fired** (the critical regression test for §19.8).
- **Probe:** `ws_probe.py` gains an `--http` mode (health, post, SSE read with
assertion flags, per `gateway-plugin/tests/README.md`).
- **Kotlin** (`:shared` commonTest): SSE parser (multi-line data, comments,
`Last-Event-ID` bookkeeping); transport state machine transitions (fake
clock: WS-loss → immediate fallback; startup race → fallback in < 1 s).
- **E2E** (`e2e.py`, new scenario): point the app at a dead WS port with the
HTTP leg live → send a message → assert user echo + agent reply arrive via
SSE; timing assertion: send → user echo < 1 s on LAN. Live-verify on the
device via ADB (screenshot of the "connected · http" pill).
## 19.13 Non-goals (v1) / future
- **Media over HTTP** (v2): `POST /v1/media` (chunked, same sha256 contract as
`07-media.md`) + `GET /v1/media/{id}` for pull/playback. Unblocks
attachments in fallback mode.
- **App-side send outbox** (companion work, separate doc): queue sends locally
when *both* legs are down; drains over whichever leg recovers. This doc
removes the 2–20 s wait; the outbox removes the last "gateway was down for
30 s" data-loss case.
- **QUIC / WebTransport** if handover flakiness persists after this + the
outbox (connection migration would make the fallback rare).
- Per-device tokens (`16-open-questions.md` #3) apply to both legs identically
when implemented.
## 19.14 Effort & change list
| Slice | Files | Est. |
| --- | --- | --- |
| Gateway leg | new `gateway-plugin/http_server.py` (~450 lines); `adapter.py` hooks (start/stop, fan-out in `_broadcast_or_log` + status path, delivery counting, dispatch-origin tag); `protocol.py` unchanged | 2–3 d |
| App leg | new `app/shared/.../net/HttpGateway.kt` (SSE reader + poll); `GatewayClient.kt` state machine + startup race; `ChatScreen.kt` gate + status pill; composer media-disable in fallback | 2–3 d |
| Tests + e2e + docs | per §19.12; `frames.schema.json` unchanged (no new frame types); `09-pairing-security.md` + `13-testing.md` cross-references | 1–2 d |
Total: **~1 week**, each slice independently shippable (gateway leg is
inert until the app uses it; app leg degrades to today's behavior if the
HTTP port is closed).
+5 -2
View File
@@ -8,6 +8,7 @@ This folder is the single source of truth for *what to build and why*. Read it
top-to-bottom once, then use the numbered docs as a lookup while implementing.
> ⚠️ **READ FIRST — two hard rules**
>
> 1. **`hermes-agent/` (sibling of this folder) is a read-only research
> reference. It must NEVER be committed, pushed, or shipped.** It is
> git-ignored at the repo root. We only *install* our plugin into a live
@@ -20,7 +21,7 @@ top-to-bottom once, then use the numbered docs as a lookup while implementing.
## Reading order
| # | File | When to read |
|---|------|--------------|
| --- | ------ | -------------- |
| 0 | [`00-overview.md`](00-overview.md) | Always first. Vision, scope, disclaimers, locked decisions. |
| 1 | [`01-architecture.md`](01-architecture.md) | Before touching code. System shape + rationale. |
| 2 | [`02-monorepo.md`](02-monorepo.md) | When scaffolding the repo. |
@@ -39,8 +40,10 @@ top-to-bottom once, then use the numbered docs as a lookup while implementing.
| 15 | [`15-hermes-reference.md`](15-hermes-reference.md) | **Cheat-sheet** of hermes-agent source to read. |
| 16 | [`16-open-questions.md`](16-open-questions.md) | Decisions made + open items. |
| 17 | [`17-future-control-surface.md`](17-future-control-surface.md) | **Backlog** — what the app could control beyond chat (cron, kanban, models, …). |
| 19 | [`19-http-fallback-transport.md`](19-http-fallback-transport.md) | **Design** — HTTP fallback leg (POST + SSE/long-poll) so the app can send/receive when the WS is down. |
Machine-readable / diagrams:
- [`protocol/frames.schema.json`](protocol/frames.schema.json) — wire-frame schema.
- [`diagrams/architecture.mmd`](diagrams/architecture.mmd) — mermaid architecture.
@@ -65,4 +68,4 @@ project** (`app/`) with a shared KMP module (`app/shared`).
- **Phase:** M0–M6 complete; M7 (polish + E2E + docs) in progress.
- **Owner decisions locked:** see [`16-open-questions.md`](16-open-questions.md).
- **Last updated:** 2026-08-19.
- **Last updated:** 2026-08-19.