A client that completes TCP but vanishes mid-TLS-handshake (e.g. a
phone losing its network/VPN while traveling) blocked
ssl.SSLSocket.accept() inside serve_forever forever: the gateway
stopped accepting any new device connections (the app could not
reconnect), and on the next restart httpd.shutdown() froze the whole
event loop until the shutdown watchdog killed the process (ARIA
journal 2026-09-11 / 2026-09-23).
- Move the TLS handshake out of the accept loop: it now runs in the
per-connection thread under a hard timeout (HANDSHAKE_TIMEOUT_S,
10 s); a failed/timed-out handshake just closes the socket.
- stop() no longer blocks the event loop: shutdown()/server_close()/
join run in an executor under asyncio.wait_for(10 s); if the bound
expires the daemon threads are abandoned.
- Regression test: a silent half-open TCP connection must not stop
fresh TLS connections from being served, and stop() must stay
bounded.
acquire_scoped_lock returns (acquired, existing_record); the old
'if not acquire_scoped_lock(...)' tested the tuple, which is always
truthy, so the 'port in use by another profile' pre-check never fired
and a conflict surfaced as a generic bind failure. Unpack and test the
first element, matching gateway/platforms/base.py's canonical usage.
Bump VERSION / plugin.yaml to 0.1.3.
- VERSION at repo root (0.1.2); bump it to cut a release
- App: generated AppVersion.kt (config-cache-safe Gradle task with
VERSION as declared input) shown in Settings; sent to the gateway
via X-Iris-App-Version header on the SSE open
- Gateway: reports its own version in hello.ack server_caps.app_version
(read from the repo-root VERSION via the plugin symlink); stores the
app's version in the device registry caps (merge, not overwrite, so
an old app reconnecting without the header doesn't wipe it)
- Settings: app + gateway version rows, mismatch hint, and a best-effort
Gitea latest-release check (ReleaseCheck) with an 'update available' hint
- Release workflow: reads VERSION from the repo (no manual input), with
a guard against an empty file
- Docs: frames.schema.json + 04-wire-protocol.md updated for app_version
adapter.py was a 3,493-line monolith. Split it into focused modules with
clear separation of responsibilities, bringing it down to ~857 lines:
- Module-level helpers: hooks, classify, pickers, commands, setup,
defaults, secrets
- Frame-handler mixins: inbound, tool_frames, push_frames, media_frames,
picker_frames, channel_frames, query_frames
- mixin_base: IrisAdapterBase (declaration-only base for shared attrs)
- adapter.py now holds only IrisAdapter (the composition of the 7 mixins
+ BasePlatformAdapter), register(), and test-facing re-exports
The mixins come before BasePlatformAdapter in the MRO so their methods
override the base; super() calls (e.g. send_image) still resolve to
BasePlatformAdapter. No circular imports; dispatch.py and http_server.py
(instance-method callers) are unaffected.
Ruff complexity ceilings (PLR0911/0912/0913/0915) restored to Ruff's
built-in defaults (12/50/6/5) instead of "just above the current maxima",
which ratchets the bar down as code grows. The existing genuinely-complex
functions (frame builders mirroring the wire schema, the QR matrix builder,
the dispatch table) carry an explicit `# noqa: PLR09xx` marking them as
reviewed, frozen exceptions; new code is held to the default ceilings.
All 125 tests green (94 test_android + 31 test_android_http); no new ruff
errors introduced.
Auth previously used the shared IRIS_TOKEN as the security principal:
a leaked token meant access to all devices, and a compromised device
could not be isolated.
Gateway:
- pairing.py: devices.token column (in-place migration) + revoked
denylist table; issue_token (idempotent, 64 hex), token_for,
reissue_token, revoke/unrevoke/is_revoked/list_revoked. The token
never leaks into device dicts (push fan-out / listings).
- http_server.py: auth accepts the shared token (bootstrap/legacy) OR
the device's own token (both constant-time); a revoked device_id is
rejected with 401 before either comparison. On SSE open (pairing)
the per-device token is minted and returned in hello.ack.
- protocol.py: hello_ack(..., device_token).
- adapter.py: setup flow (hermes gateway setup -> Iris) now offers
'Remove a paired device?' on an existing setup: numbered select
menu (last option = exit the removal loop), confirmation, back to
the menu for further removals.
- tools/iris_devices.py: operator CLI (list / revoke / unrevoke /
reissue), stdlib only.
App:
- SecureStore.deviceToken (Android: EncryptedSharedPreferences;
Desktop: second keyring slot iris-device-token / device_token.enc).
- HelloAckPayload.deviceToken; GatewayClient stores it on hello and
presents it instead of the shared token from then on (live provider
in HttpGateway); savePairing/clear wipe it for re-pairing.
Docs: 09 §9.3 stretch -> implemented (revocation semantics, both
control surfaces), 04 hello.ack example, frames.schema.json, M7 row 13.
Tests: 8 new Python tests (issuance, acceptance, revocation,
isolation, unrevoke, registry unit x2, setup-flow menu) - 94/94 pass;
2 new Kotlin wire tests - green. Live-verified against a running
gateway (hello.ack token matches devices.db; revoke -> 401 even with
shared token; unrevoke -> 200; setup TUI both paths).
- outbox: delete_message/message_info now match the exact lane first
(a flat-lane delete/lookup with thread_id=None sees only frames with
no thread_id) and fall back to the message_id across all lanes only
when the exact lane matches nothing. Previously lane=None meant
'any lane' in the first pass, so a flat-lane delete also removed
same-id frames from threads (test expected 3 removed, got 4).
- http_server: the SSE live loop skipped queued frames when stop() set
sub.closed before the handler thread reached the loop (descheduled
under load between the initial hello/status writes and the loop).
The loop now drains frames queued before the close, so the
status{restarting} teardown broadcast always reaches the client
before EOF (test_disconnect_broadcasts_status_restarting was flaky
~70% under CPU load).
Add a todo.update frame (server->app) carrying the agent's full current
todo list. The gateway emits it whenever the hermes todo tool completes
(the tool result is authoritative even for merge writes) and re-sends a
snapshot right after hello so a reconnecting device re-learns the plan.
Ephemeral: never outboxed.
The app renders it as a compact strip above the composer (max 3 lines,
the rest scrollable) mirroring the hermes desktop composer status stack:
pending = hollow ring, in_progress = spinner, completed = green check,
cancelled = struck through. It auto-scrolls to the current task whenever
the active task changes, and hides itself once the list is empty or fully
resolved.
Gateway (docs/19):
- Remove ws_server.py; frame dispatch factored into dispatch.py
- http_server: media upload/pull, pairing over HTTP
- protocol: media frames mirrored; tests + ws_probe updated for HTTP
App:
- HttpGateway: postFrame/uploadMedia/pullMedia no longer throw on
network failure (PostResult ok=false / Result.failure) — uncaught
SocketTimeoutException on Dispatchers.Default crashed the app
- GatewayClient: dead-stream watchdog (health probe every 10s, 2
failures -> redial in ~20s instead of the 45s SSE read timeout);
state flips to Reconnecting when the stream dies, restored from the
last hello.ack on long-poll success; poke() + backoff reset on app
resume (MainActivity.onResume)
- Offline sends: composer enabled while disconnected; a send with no
response (status 0) stays queued (Pending) and is auto-resent on the
next (re)connect after a 2s outbox-replay grace; gateway 4xx
rejections fail the bubble (tap to retry, no auto-loop)
- ChatStore: echo-replace and thread-relocate also match Failed
bubbles (POST response lost in a network drop); loadHistory dedupes
local failed bubbles the server already has; failMessage()
- MainActivity: poke() on resume so a backgrounded app reconnects
promptly instead of waiting out the backoff
Second, short-lived-connection transport next to the WS: same frames,
same outbox/cursor, same token, served over plain HTTP (stdlib
ThreadingHTTPServer bridged into the asyncio loop; zero new deps).
- http_server.py: /v1/health (unauthenticated), POST /v1/frame
(accept-and-ack; validation rejections as 4xx error frames), SSE
/v1/events (outbox catch-up with id=cursor, event: hello, 15s
heartbeat, bounded-queue backpressure), long-poll /v1/poll (25s hold).
Bearer token + X-Iris-Device (same allowlist as WS hello), 64 KiB body
cap, per-device rate limit, optional TLS, non-fatal bind failure.
- ws_server.py: inbound dispatch chain extracted to shared
dispatch_frame() used by both transports.
- adapter.py: ANDROID_HTTP_PORT/CERT/KEY config; start/stop next to the
WS; delivery counting in _broadcast_or_log (an SSE subscriber is a
live subscriber -> no push, docs/19 19.8); _reply() routes
point-to-point replies into the in-flight HTTP response (reply sink)
or broadcasts when the device has no live WS (19.7); status/typing/
channel events fan out to both transports.
- ws_probe.py: --http mode (health + POST + SSE turn drive, same
assertion flags); tests/README updated.
- Tests: hermes-agent/tests/gateway/test_android_http.py (23 tests,
incl. the 19.8 delivery-counting regression); test_android.py (74)
still green.