hermes gateway setup now asks 'Set up TLS now?' when IRIS_HTTP_CERT is
not in .env (default No, Yes for an all-interfaces bind). Accepting
generates a 10-year RSA-2048 self-signed cert with SANs (advertised LAN
IP, hostname, loopback) under ~/.hermes/iris/ via hermes' existing
cryptography dependency, saves IRIS_HTTP_CERT/IRIS_HTTP_KEY, and prints
the SHA-256 fingerprint in openssl format for the app's confirm-and-pin
dialog. The pairing URL/QR printed afterwards already advertise https.
- key created 0600 from the start (no umask window)
- save_env_value inside the best-effort guard (unwritable .env warns)
- leftover cert without env var -> overwrite confirmation (protects the
app's pinned fingerprint)
- bind wildcards (0.0.0.0 / ::) never become SANs; :: gets the same
default-Yes as 0.0.0.0 (pairing._unroutable parity)
Tests: 5 new (cert generation incl. openssl fingerprint cross-check,
accept/decline, no re-prompt, default-follows-bind, overwrite prompt).
Docs: install.md Part 2 table + Part 4 Option B.
- docs/install.md: new end-to-end guide for non-technical users
(gateway install, app install, LAN/TLS/remote connection, push,
options, troubleshooting); docs/setup.md now points to it
- README: new 'Install the gateway' section; pairing section updated
for HTTP transport (8791, QR scan on Android)
- rename IRIS_WS_HOST -> IRIS_HTTP_HOST (clean rename, no compat
fallback); drop dead DEFAULT_PORT=8790
- setup.py: advertise https:// in the printed/QR server URL when
IRIS_HTTP_CERT is set
- ws_probe.py/e2e.py: default --url http://127.0.0.1:8791, env
IRIS_WS_URL -> IRIS_HTTP_URL, honor explicit port + https scheme
- plugin.yaml: IRIS_HTTP_* env names, description no longer says
'WebSocket server'
- docs 03/09/12/19: fix stale WS-era refs (ws_server.py cites,
8790 smoke test, WSS->HTTPS, 'HTTP fallback' reframed as the
only transport)
- AGENTS.md: symlink name android -> iris (matches actual install)
- test: adapter reads IRIS_HTTP_HOST/CERT/KEY from env; legacy
IRIS_WS_* names are not consulted (95/95 pass)
The plugin is named 'iris' (IrisAdapter, IRIS_HOME_CHANNEL, label Iris),
but several docs still referred to it as the android platform/plugin and
to the product as 'the Android app'. Rename name-mentions to iris/IRIS
and product-mentions to 'Iris app'; keep legitimate OS references
(androidApp, Android SDK, Android 10, androidx, test_android.py, ...).
Also includes pi-lens markdown-lint autofixes (table spacing, trailing
newlines) in the touched files.
docs/09 §9.4 promised a SSH-host-key-style fingerprint confirm on first
pair, but the app built a default OkHttpClient with no certificate
handling — self-signed gateway certs were simply rejected.
Build the flow:
- TlsPinning.kt: PinningTrustManager wraps the platform default trust
manager; a rejected cert is accepted only when its SHA-256 fingerprint
matches the user-confirmed pin, anything else fails with
TlsFingerprintRequired (hostname verification still applies).
- SecureStore.pinnedCertFingerprint (Android EncryptedSharedPreferences +
desktop settings.json), cleared on forget().
- GatewayClient: pinning socket factory on the shared client (all legs
inherit it), new terminal State.TlsConfirmRequired, unwrap the nested
TlsFingerprintRequired in connect loop / watchdog / SSE / poll /
testHello.
- ConnectScreen: confirm dialog showing the fingerprint ("Confirm &
pin" re-runs the connect); IrisApp routes TlsConfirmRequired there;
ChatScreen + desktop tray handle the new state.
- Docs: §9.4 now describes the real flow (incl. SAN requirement), gap
table item 6 → implemented, setup.md limitation note updated.
- Tests: TlsPinningTest (fingerprint vs openssl, pin accept/reject,
live pin read, unwrap) + TlsPinningIntegrationTest (real TLS
handshake: unpinned → confirm data, pinned → 200).
Live E2E verified on the phone: first pair against a self-signed
IRIS_HTTP_CERT gateway shows the dialog, confirm pins, chat works,
auto-reconnect after gateway restart uses the pin.
- 'No file limit' -> 100 MB default, configurable via max_upload_bytes
- 'No character limit' -> 'No 4,096-character limit like Telegram'
(hard frame-body cap: 1 MiB)
- Note that limits are set on the gateway side (hermes), not in the app
- Sync docs/playstore-listing.md (same overclaim)
Auth previously used the shared IRIS_TOKEN as the security principal:
a leaked token meant access to all devices, and a compromised device
could not be isolated.
Gateway:
- pairing.py: devices.token column (in-place migration) + revoked
denylist table; issue_token (idempotent, 64 hex), token_for,
reissue_token, revoke/unrevoke/is_revoked/list_revoked. The token
never leaks into device dicts (push fan-out / listings).
- http_server.py: auth accepts the shared token (bootstrap/legacy) OR
the device's own token (both constant-time); a revoked device_id is
rejected with 401 before either comparison. On SSE open (pairing)
the per-device token is minted and returned in hello.ack.
- protocol.py: hello_ack(..., device_token).
- adapter.py: setup flow (hermes gateway setup -> Iris) now offers
'Remove a paired device?' on an existing setup: numbered select
menu (last option = exit the removal loop), confirmation, back to
the menu for further removals.
- tools/iris_devices.py: operator CLI (list / revoke / unrevoke /
reissue), stdlib only.
App:
- SecureStore.deviceToken (Android: EncryptedSharedPreferences;
Desktop: second keyring slot iris-device-token / device_token.enc).
- HelloAckPayload.deviceToken; GatewayClient stores it on hello and
presents it instead of the shared token from then on (live provider
in HttpGateway); savePairing/clear wipe it for re-pairing.
Docs: 09 §9.3 stretch -> implemented (revocation semantics, both
control surfaces), 04 hello.ack example, frames.schema.json, M7 row 13.
Tests: 8 new Python tests (issuance, acceptance, revocation,
isolation, unrevoke, registry unit x2, setup-flow menu) - 94/94 pass;
2 new Kotlin wire tests - green. Live-verified against a running
gateway (hello.ack token matches devices.db; revoke -> 401 even with
shared token; unrevoke -> 200; setup TUI both paths).
- IRIS_PUSH_BACKEND now defaults to ntfy (keeps push metadata on your own
infrastructure); FCM is opt-in via IRIS_PUSH_BACKEND=fcm
- build_push_backend(): ntfy for empty/unknown names, FCM only on explicit 'fcm'
- gateway setup: warn when FCM is chosen (metadata routed via Google's servers)
- README: privacy note + dedicated push section; new docs/playstore-listing.md
with the FCM/ntfy privacy note for the Play Store listing
- docs: 00/02/03/08/12/16 + setup.md updated to ntfy-default wording
- tests: default-backend assertion updated (86/86 pass)
Add a todo.update frame (server->app) carrying the agent's full current
todo list. The gateway emits it whenever the hermes todo tool completes
(the tool result is authoritative even for merge writes) and re-sends a
snapshot right after hello so a reconnecting device re-learns the plan.
Ephemeral: never outboxed.
The app renders it as a compact strip above the composer (max 3 lines,
the rest scrollable) mirroring the hermes desktop composer status stack:
pending = hollow ring, in_progress = spinner, completed = green check,
cancelled = struck through. It auto-scrolls to the current task whenever
the active task changes, and hides itself once the list is empty or fully
resolved.
Slash commands with a finite set of options (/reasoning, /fast, ...) now
render a tappable card with buttons (2 per row, ✓ on the current value)
instead of a plain text status card. The mechanism is generic: any command
that calls the adapter's send_choice_picker() gets a picker automatically.
Wire protocol (docs/04, frames.schema.json):
- picker.choice (server→app): {picker_id, title, choices[]}
- picker.select (app→server): {picker_id, value}
- pickers capability flag now True in server_caps
gateway-plugin:
- protocol.py: picker.choice/picker.select frame types + picker_choice()
- dispatch.py: route picker.select → adapter.on_picker_select
- adapter.py: send_choice_picker() (fails cleanly with no live device so
hermes falls back to text), on_picker_select(), in-memory pending pickers
(gateway restart expires them; stale select is a no-op), pickers=True
app (KMP):
- Protocol.kt: PickerChoice/PickerChoicePayload + pickerSelectFrame()
- ChatStore.kt: PickerItem + onPickerChoice (idempotent) + resolvePicker
(optimistic, one-shot)
- ChatDb.kt: persist PickerItem in the messages table (polymorphic decode)
- IrisController.kt: picker.choice routing + selectPicker() action
- ChatScreen.kt: PickerCard composable (locks after selection)
Tests:
- python: 3 picker tests (roundtrip, no-device fallback, stale-select noop)
- kotlin: ChatStorePickerTest (add/idempotent/resolve/one-shot/noop/serialize)
- fixture fix: clear leaked IRIS_HTTP_PORT/IRIS_WS_HOST env so the adapter
binds the ephemeral port (a prior test's interactive_setup() polluted the
process env, colliding with a live gateway on 8791)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
tool.start gains an optional cosmetic 'emoji' field, resolved server-side
through hermes' own display layer (active-skin overrides, then the tool
registry's per-tool emoji) so icons track the user's hermes theme and
new/plugin tools get their glyph for free. Omitted for unknown tools so
the app falls back to its default wrench.
- protocol.py: tool_start(emoji=...) kwarg, payload field when set
- adapter.py: _tool_emoji() helper (lazy import, None on unknown/failure)
- frames.schema.json + docs/04: field documented
- app: ToolStartPayload.emoji -> ToolItem.emoji -> ToolCard header
- tests: frame shape, resolution/fallback, end-to-end tool.start emoji
Gateway (docs/19):
- Remove ws_server.py; frame dispatch factored into dispatch.py
- http_server: media upload/pull, pairing over HTTP
- protocol: media frames mirrored; tests + ws_probe updated for HTTP
App:
- HttpGateway: postFrame/uploadMedia/pullMedia no longer throw on
network failure (PostResult ok=false / Result.failure) — uncaught
SocketTimeoutException on Dispatchers.Default crashed the app
- GatewayClient: dead-stream watchdog (health probe every 10s, 2
failures -> redial in ~20s instead of the 45s SSE read timeout);
state flips to Reconnecting when the stream dies, restored from the
last hello.ack on long-poll success; poke() + backoff reset on app
resume (MainActivity.onResume)
- Offline sends: composer enabled while disconnected; a send with no
response (status 0) stays queued (Pending) and is auto-resent on the
next (re)connect after a 2s outbox-replay grace; gateway 4xx
rejections fail the bubble (tap to retry, no auto-loop)
- ChatStore: echo-replace and thread-relocate also match Failed
bubbles (POST response lost in a network drop); loadHistory dedupes
local failed bubbles the server already has; failMessage()
- MainActivity: poke() on resume so a backgrounded app reconnects
promptly instead of waiting out the backoff
- e2e.py: scenario 13 (http fallback) — drives a full turn over the
HTTP leg (health + POST /v1/frame + SSE /v1/events, no WS) and
asserts the user echo lands on the SSE stream in < 1.5 s.
- ws_probe.py --http: prints '== user echo in X.XXs' (the docs/19
sendable-in-fallback timing assertion) alongside the existing
health/POST/SSE output; same assertion flags as the WS leg.
- docs: 19 status flipped to implemented; 09-pairing-security §9.4
cross-reference (second door, same lock: token + device allowlist,
64 KiB cap, rate limit, optional TLS, unauthenticated /v1/health);
13-testing manual scenario 15 + automated pointers.
Second, short-lived-connection transport next to the WS: same frames,
same outbox/cursor, same token, served over plain HTTP (stdlib
ThreadingHTTPServer bridged into the asyncio loop; zero new deps).
- http_server.py: /v1/health (unauthenticated), POST /v1/frame
(accept-and-ack; validation rejections as 4xx error frames), SSE
/v1/events (outbox catch-up with id=cursor, event: hello, 15s
heartbeat, bounded-queue backpressure), long-poll /v1/poll (25s hold).
Bearer token + X-Iris-Device (same allowlist as WS hello), 64 KiB body
cap, per-device rate limit, optional TLS, non-fatal bind failure.
- ws_server.py: inbound dispatch chain extracted to shared
dispatch_frame() used by both transports.
- adapter.py: ANDROID_HTTP_PORT/CERT/KEY config; start/stop next to the
WS; delivery counting in _broadcast_or_log (an SSE subscriber is a
live subscriber -> no push, docs/19 19.8); _reply() routes
point-to-point replies into the in-flight HTTP response (reply sink)
or broadcasts when the device has no live WS (19.7); status/typing/
channel events fan out to both transports.
- ws_probe.py: --http mode (health + POST + SSE turn drive, same
assertion flags); tests/README updated.
- Tests: hermes-agent/tests/gateway/test_android_http.py (23 tests,
incl. the 19.8 delivery-counting regression); test_android.py (74)
still green.
The app previously showed 'Gateway restarting' on every connection loss.
Now the gateway broadcasts status{state=restarting} on its shutdown path
(before closing the sockets), and the app:
- posts 'Gateway restarting' immediately on that frame (not on the
socket-drop transition, which lags by the ~20s WS ping timeout)
- posts 'Gateway online' on the next reconnect only when the restart
notice was posted (latch) - a plain network drop shows neither, just
the reconnect banner
- drops the 'Gateway is restarting...' banner (replaced by the chat notice)
Docs (04-wire-protocol, frames.schema.json) updated: restarting is no
longer reserved. Test for the disconnect broadcast added to the local
hermes-agent test mirror (git-ignored, not committed).
- Cache.sq / ChatDb: lanes, channels and last_lane persisted as JSON
snapshots; sanitize on restore (Pending->Failed, streaming->false);
ephemeral items (tool, system) skipped
- Platform drivers via expect/actual: AndroidSqliteDriver (app db dir)
/ JdbcSqliteDriver (~/.iris/iris_cache.db)
- IrisController: restore before connect, debounced (750ms) snapshot
persistence, synchronous flush on dispose, clearAll on forget
- History robustness: historyLoaded marked only when the response is
processed (lost request/response retried on reconnect); events
collector wrapped in try/catch; onHelloAck fast path so the history
request fires on the WS thread instead of the starved state
collector; frame-decode and history-load logging
- Protocol: HistoryMessage.media nullable (defensive vs older
gateways that sent "media": null)
- Tests: ChatDbTest (jvmTest, JDBC in-memory), ChatStoreCacheTest,
HistoryPayloadTest, HistoryWireTest (real captured 92KB response)
- docs/10: §10.7 implemented schema, new §10.9 cache behavior
Deleting a thread, channel, or message was a no-op/soft-delete: messages
were only dropped from the plugin outbox (still in hermes' session store,
hence searchable/recoverable) and channels/threads were merely archived.
Now deletion is complete and non-recoverable, with no search trace:
- purge.py (new): hard-delete from hermes' session store (state.db).
delete_lane wipes a channel's/thread's sessions + messages; deleting a
messages row also drops it from the FTS5 index via the delete triggers.
delete_message removes one message, matched by (session, role, exact
content, closest timestamp) since plugin m_<hex> ids aren't persisted.
- channels.py: delete() hard-deletes the row (and a channel's child
threads) instead of archiving.
- outbox.py: add delete_lane() (wipe all frames for a lane) and
message_info() (read a message's final role/text/ts for the match).
- adapter.py: on_channel_delete wipes outbox + session store;
on_message_delete purges the session-store row per message.
- App: delete confirmations no longer claim history stays for search;
ChannelStore removes a channel's threads on channel delete.
- Docs updated to describe hard deletion.
Hermes can append a text "runtime footer" (model, context %, workdir,
latency, cost) to final replies, but only when display.runtime_footer is
enabled in the hermes config. We want the same info but controlled by the
APP, not the gateway config. So the gateway now ALWAYS sends the data as a
structured `runtime` object on final assistant messages, and the app decides
whether/what to show.
Gateway (gateway-plugin/):
- protocol.py: new runtime_footer() helper + `runtime` field on the
message / message.stop frames. Keys (all optional, absent when the data
is unavailable — e.g. no cost for local models): model (vendor prefix
dropped), context_pct (0-100), cwd (home-relative), latency (seconds),
cost (USD).
- adapter.py: a post_api_request plugin hook captures the turn's model +
prompt tokens + start time (platform-filtered to android so other
platforms don't pollute the buffer). _build_runtime_footer() resolves the
model's context window (cached, best-effort, off the event loop via
asyncio.to_thread with a timeout) and computes context_pct. The runtime
object is attached on every final send (streaming message.stop and
non-streaming message, plus the fallback paths).
- outbox.py: `runtime` preserved in history reconstruction so the footer
survives a restart / first open.
App (app/shared/):
- Protocol.kt: RuntimeMeta data class + `runtime` on MessagePayload /
MessageStopPayload / HistoryMessage.
- ChatStore.kt: `runtime` on MessageItem, wired through live + history
reconciliation.
- SecureStore.kt (+ Android/Desktop actuals): runtimeFooterEnabled +
runtimeFooterFields (persisted per device).
- IrisController.kt: StateFlows + toggleRuntimeFooter() /
toggleRuntimeField(); RUNTIME_FIELD_KEYS / default set / parser.
- SettingsScreen.kt: "Runtime footer" switch; when on, an expandable chip
menu (Model · Context % · Workdir · Latency · Cost) to pick fields.
- ChatScreen.kt: footer rendered on the SAME line as the timestamp (footer
left, time right, Telegram-style), only for final non-streaming assistant
answers; Inspector pane now shows the runtime fields too.
Docs: 04-wire-protocol.md + frames.schema.json document the `runtime`
object.
Verified end-to-end on device: final replies carry
`qwen3.8-27B-exl3-4.5bpw · 53% · ~ · 38s` with the time right-aligned on
the same line; 69/69 gateway tests pass, Kotlin builds + tests pass.
- Auto-threading (Telegram topic-mode workflow): message.send {auto_thread}
mints a fresh AI-named thread (instant derived title, LLM upgrade via
channel.renamed); channel.created {auto:true}; the app jumps into the new
thread and relocates the optimistic pending bubble.
- history frame: paged full message history for initial channel open /
scroll-up pagination (reconstructed from the outbox log).
- Streaming on/off: gateway side (display.platforms.android.streaming) plus a
per-device app toggle (Settings → Streaming); reasoning/model/tokens carried
on message frames.
- Context menu: long-press (touch) / right-click (desktop) thread affordances
via a KMP rightClick expect/actual.
- Settings → Reasoning: auto-collapse long reasoning blocks (default on).
- Tool detail: the gateway now always supplies full tool data — it forces
verbose tool progress (full args → tool.start.args) and captures each
completed call via the post_tool_call hook (output/duration/ok → tool.end).
The app reveals the full call + output on expand (Truncated) and
auto-expands cards in Everything mode.