- IRIS_PUSH_BACKEND now defaults to ntfy (keeps push metadata on your own infrastructure); FCM is opt-in via IRIS_PUSH_BACKEND=fcm - build_push_backend(): ntfy for empty/unknown names, FCM only on explicit 'fcm' - gateway setup: warn when FCM is chosen (metadata routed via Google's servers) - README: privacy note + dedicated push section; new docs/playstore-listing.md with the FCM/ntfy privacy note for the Play Store listing - docs: 00/02/03/08/12/16 + setup.md updated to ntfy-default wording - tests: default-backend assertion updated (86/86 pass)
13 KiB
03 — Gateway Plugin (Python)
The plugin is a community-style hermes platform plugin named iris.
It follows the "Plugin Path" in hermes-agent/gateway/platforms/ADDING_A_PLATFORM.md
and the canonical example hermes-agent/plugins/platforms/irc/adapter.py.
Zero hermes-core changes. Zero new Python dependencies (websockets and
httpx are already core deps).
Source map for every hermes integration point: see
15-hermes-reference.md.
3.1 plugin.yaml (manifest)
name: iris-platform
label: Android
kind: platform
version: 0.1.0
description: >
Native Android / Desktop client gateway adapter for Hermes Agent.
Runs a WebSocket server inside the gateway; the app connects with a
pairing token. Supports streaming, reasoning, structured tool events,
channels/threads, media, FTS5 search, and FCM/ntfy push.
author: <you>
requires_env:
- name: IRIS_TOKEN
description: "Shared pairing token the app presents on connect"
prompt: "Android pairing token"
password: true
optional_env:
- name: IRIS_WS_HOST
description: "WS bind host (default 127.0.0.1; use 0.0.0.0 for LAN)"
prompt: "WS host"
password: false
- name: IRIS_WS_PORT
description: "WS port (default 8790)"
prompt: "WS port"
password: false
- name: ANDROID_HOME_CHANNEL
description: "Default chat id for cron/notification delivery (default default)"
prompt: "Home channel"
password: false
- name: IRIS_ALLOWED_USERS
description: "Comma-separated allowed device_ids (empty = token-only auth)"
prompt: "Allowed device ids"
password: false
- name: IRIS_ALLOW_ALL_USERS
description: "Allow any paired device (dev only)"
prompt: "Allow all devices? (true/false)"
password: false
- name: IRIS_PUSH_BACKEND
description: "Push backend: ntfy (default, keeps metadata off Google) or fcm"
prompt: "Push backend"
password: false
- name: IRIS_FCM_SERVICE_ACCOUNT
description: "Path to Firebase service-account JSON (FCM HTTP v1)"
prompt: "FCM service account path"
password: true
- name: IRIS_FCM_SERVER_KEY
description: "Legacy FCM server key (fallback if no service account)"
prompt: "FCM server key"
password: true
- name: NTFY_TOPIC
description: "ntfy topic for push (when IRIS_PUSH_BACKEND=ntfy)"
prompt: "ntfy topic"
password: false
- name: NTFY_SERVER_URL
description: "ntfy server URL (default https://ntfy.sh)"
prompt: "ntfy server URL"
password: false
- name: IRIS_WS_CERT
description: "TLS cert path for WSS (optional)"
prompt: "WSS cert"
password: false
- name: IRIS_WS_KEY
description: "TLS key path for WSS (optional)"
prompt: "WSS key"
password: false
Behavioral (non-secret) settings live in config.yaml under
gateway.platforms.iris.extra (host, port, home_channel, outbox retention,
max upload bytes, tls). Secrets live in .env. (hermes policy: .env = secrets
only.)
3.2 register(ctx) entry point
def register(ctx):
ctx.register_platform(
name="iris",
label="Iris",
adapter_factory=lambda cfg: IrisAdapter(cfg),
check_fn=check_requirements, # passive: websockets importable + token set
validate_config=validate_config, # host/port/token present
is_connected=is_connected,
required_env=["IRIS_TOKEN"],
install_hint="No extra packages needed (websockets + httpx are core deps)",
setup_fn=interactive_setup, # hermes gateway setup flow
env_enablement_fn=_env_enablement, # seed extra + home_channel from env
cron_deliver_env_var="ANDROID_HOME_CHANNEL",
standalone_sender_fn=_standalone_send, # best-effort out-of-proc cron (stretch)
parse_target_ref_fn=_parse_target_ref, # "iris:<chat>[:<thread>]"
allowed_users_env="IRIS_ALLOWED_USERS",
allow_all_env="IRIS_ALLOW_ALL_USERS",
max_message_length=0, # 0 = no limit (WS has none)
emoji="📱",
pii_safe=False,
platform_hint=(
"You are chatting with the user through their native Iris app "
"(Android/Desktop). It renders Markdown, inline code, images, "
"audio and video, and shows your reasoning and tool activity. "
"Conversations are organized into channels and optional threads. "
"Keep formatting rich but readable."
),
)
Field reference (all from PlatformEntry, gateway/platform_registry.py:63):
adapter_factory, check_fn, validate_config, is_connected, required_env,
install_hint, setup_fn, env_enablement_fn, apply_yaml_config_fn,
cron_deliver_env_var, parse_target_ref_fn, allowed_users_env,
allow_all_env, max_message_length, pii_safe, platform_hint, emoji,
ensure_deps_fn.
check_requirements()— passive probe:import websocketssucceeds andIRIS_TOKENis set. Never installs._env_enablement()— returns a dict seedingPlatformConfig.extra(host/port/home_channel/push_backend) + ahome_channelkey{"chat_id": "default", "name": "Default"}sohermes gateway statusand cron home-channel resolution work without instantiating the adapter._parse_target_ref(ref)— the core strips the platform prefix first, sorefis the direct chat id (e.g.chan_7,default) with an optional:t_<n>thread suffix; friendly names resolve via the channel directory. Returns(chat_id, thread_id)orNone.interactive_setup()— prompts for token (or generates one), host/port, push backend + credentials, prints a QR code (pairing) and the app URL.
3.3 IrisAdapter(BasePlatformAdapter)
Constructor: super().__init__(config=config, platform=Platform("iris")).
Reads config.extra (env overrides win). Initializes: WS server (not started
until connect()), connection registry, outbox (SQLite under
get_hermes_home()/"iris"), push backend, pairing store, channel directory.
Lifecycle
connect(*, is_reconnect=False) -> bool- Acquire scoped lock (
gateway.status.acquire_scoped_lock("iris", key)) so two profiles can't bind the same port/identity. - Start the
websocketsserver onhost:port(TLS if cert/key set). _mark_connected(); return True.
- Acquire scoped lock (
disconnect()- Stop server, close all device sockets, release lock,
_mark_disconnected().
- Stop server, close all device sockets, release lock,
Inbound (app → agent)
- WS
message.send {text, reply_to?, media_refs?}→ buildSessionSourceviaself.build_source(chat_id, chat_name, chat_type, user_id, user_name, thread_id)→ buildMessageEvent(text=…, message_type=TEXT, source=…, media_urls=[cached paths], media_types=[…], reply_to_message_id=…)→await self.handle_message(event). - Slash commands arrive as plain text starting with
/; the gateway's command pipeline resolves + dispatches them (no special handling needed). picker.select {picker_id, value}→ route to the gateway-side resolvers (model picker / choice picker / clarify / approval / slash-confirm) using the shared callback-id conventions (cl:<id>:<idx>,appr:<id>:<choice>,sc:<choice>:<id>).channel.create/channel.rename/channel.set_default→ mutate the channel directory (SQLite) + emitchannel.*frames to all devices.search {query, scope, chat_id?, thread_id?}→search.py→search.results.media.upload(chunked) →media.py→cache_*_from_bytes→media_ref.media.pull {media_id}→ stream cached bytes as binary frames.fcm.register {token}/hello→ update device registry.read.receipt {message_id}→ mark delivered/read (drives ✓✓), ack.sync {cursor}→outbox.py→ replay frames since cursor.
Outbound (agent → app)
send(chat_id, content, reply_to=None, metadata=None) -> SendResult- Split reasoning prefix (see
05-streaming.md) →reasoningfield. - If any device is connected: broadcast
messageframe to all. - Else (no live devices): append to outbox + fire push (FCM/ntfy).
- Return
SendResult(success=True, message_id=<id>).
- Split reasoning prefix (see
edit_message(chat_id, message_id, content)→message.updateframe (drives streaming). If no device, no-op (outbox holds the finalsend).send_typing(chat_id, metadata=None)→typingframe.get_chat_info(chat_id) -> dict→{"name": <channel name>, "type": "dm"|"channel"}from the channel directory.- Media send —
send_image / send_video / send_document / send_voice / send_image_file / send_multiple_images: stage the file in the media cache, mint amedia_id, emitmedia.offer {media_id, mime, size, filename, kind}, serve bytes onmedia.pull. (Base-classextract_media/extract_imagesalready pullMEDIA:/image tags out of agent text and call these.) - Interactive pickers —
send_model_picker(...),send_choice_picker(...),send_clarify(...),send_exec_approval(...),send_slash_confirm(...): emitpicker.model/picker.choice/picker.clarify/picker.approval/picker.confirmframes with options; store pending state keyed bypicker_id; resolve onpicker.select. create_handoff_thread(chat_id, name)→ create a thread id, register in channel directory, return it (used by cron "continuable" threads).
Streaming hooks
The main gateway drives delivery through the legacy callback path:
stream_delta_callback→GatewayStreamConsumer→send()(first) +edit_message()(updates) →message.start/message.update.tool_progress_callback→ progress queue →send_progress_messages→send()→tool.*frames (structured; classified in the adapter).interim_assistant_callback→ consumeron_commentary→send()→commentaryframes.
The adapter tracks per-chat turn state (in-turn, current streaming
message_id, last tool index) to classify outbound send() calls into
message vs tool.* vs commentary. The exact classification markers are
verified empirically in M2 (see 13-testing.md).
3.4 WebSocket server (ws_server.py)
- Library:
websockets(core dep, v15).websockets.serve(handler, host, port, ssl=ctx). - Handler per connection:
- Await first frame; must be
hello {token, device_id, device_name, caps, fcm_token?}. Verify token (constant-time) + allowlist. On failure: senderror {code:"auth"}and close. - On success: register in connection registry
(
device_id → {ws, caps, fcm_token}), sendhello.ack {server_caps, sync_cursor, channels[]}. - Loop: decode frames, dispatch to adapter inbound handlers.
- On close: deregister; if no devices remain, ensure pending outbox frames have push fired.
- Await first frame; must be
- Routing:
emit(chat_id, frame)→ broadcast to all connected devices (no per-chat subscribe; single-user model). Global frames (channel.*,status) also broadcast to all. - Heartbeat: WS ping/pong + app-level
ping/pong; dead peers reaped. - Backpressure: per-connection send queue with a bounded buffer; drop
message.update(coalesce to latest) under pressure, never dropmessage/tool.end/notification.
3.5 State & storage (all under get_hermes_home()/"iris")
Use
get_hermes_home()fromhermes_constantsfor all paths (profile-safe). Never hardcode~/.hermes.
devices.db— device registry (device_id, name, caps, fcm_token, ntfy_topic, last_seen, created).channels.db— channel directory (chat_id, name, kind: default|channel|thread, parent_chat_id, created, is_default).outbox.db— undelivered frames per chat_id + monotonic cursor.media/— inbound + outbound media cache (reuse hermescache_*_from_bytesdirs where possible).
3.6 Config resolution
- Secrets (
.env):IRIS_TOKEN,IRIS_FCM_SERVICE_ACCOUNT,IRIS_FCM_SERVER_KEY,IRIS_WS_CERT/KEY,NTFY_TOPIC(if secret). - Behavioral (
config.yaml→gateway.platforms.iris.extra):host,port,home_channel,allowed_users,push_backend,outbox_retention_hours,max_upload_bytes,tls. - Env vars override
config.yaml(hermes convention). Read secrets with the scope-aware_get_scoped_secretpattern (seeplugins/platforms/irc/adapter.py:42) so multiplexed profiles don't leak each other's tokens.
3.7 Failure & lifecycle safety
- WS server bind failure →
_set_fatal_error("bind_failed", …, retryable=True). - All outbound sends are best-effort; a dead socket latches and the frame falls to the outbox.
disconnect()cancels the server task and closes sockets cleanly.- Token/PII redaction in all logs (hermes PII policy).