# 05 — Streaming, Reasoning, Tools, Intermediate Messages This doc covers the four "agent transparency" features and exactly how each is produced by the gateway and rendered by the app. > The main hermes gateway delivers via the **legacy callback path** (not the > ACP-only event-native `render_message_event` path). We map the legacy > `send`/`edit_message`/progress calls to our WS frames. ## 5.1 Text streaming **Gateway side.** The agent's `stream_delta_callback` feeds a `GatewayStreamConsumer` (`gateway/stream_consumer.py:156`). The consumer accumulates text and, at intervals / thresholds, calls: - `adapter.send(chat_id, text)` — first time a bubble is created. - `adapter.edit_message(chat_id, message_id, text)` — subsequent updates (each carries the **full** accumulated text). **Adapter → frames.** - First `send()` of a turn segment → `message.start {message_id, role}`. - Each `edit_message()` → `message.update {message_id, text}` (full text). - Segment/turn finalization → `message.stop {message_id, final_text, reasoning?, model?, tokens?}` (or a final `message` frame when not streaming). **App side.** Maintain a live bubble per `message_id`. On `message.update`, replace the bubble text (cheap: it's a full snapshot). On `message.stop`, finalize (attach reasoning/model/tokens footer, stop the cursor). Auto-scroll while the user is at the bottom. **Streaming on/off.** Two levels: - **Gateway side:** hermes `display.platforms.iris.streaming` (default follows global). When off, the app just gets one final `message` frame. - **App side (per device):** Settings → "Streaming" toggle (default on). When off, the app ignores `message.start`/`message.update` frames and materializes each reply as a single final message on `message.stop` (`ChatStore.onMessageStop` already handles a stop without a live bubble). The gateway keeps streaming — frames are broadcast to all devices, so this is a display preference like tool verbosity, not a wire flag. ## 5.2 Reasoning (shown *before* the message) **Requirement:** reasoning appears as a block **above** the answer (like the reference screenshot's "Reasoning:" panel with a copy button). **Gateway side.** hermes prepends reasoning to the final response when `show_reasoning` is enabled (`gateway/run.py:20089`). The format is stable and chosen by `reasoning_style` (`gateway/display_config.py:37`): - `code` (default): `💭 **Reasoning:**\n```\n\n```\n\n` - `blockquote`: `> 💭 **Reasoning:**\n> …\n\n` - `subtext`: `-# 💭 Reasoning\n-# …\n\n` (Discord-style) **Plugin config.** Set for the `iris` platform: ```yaml display: platforms: iris: show_reasoning: true reasoning_style: code # we split on the code-fence form ``` **Adapter split.** In `send()`, detect the `code`-style prefix and split: ``` prefix = "💭 **Reasoning:**\n```\n" # find the closing "\n```\n\n" after the prefix reasoning = text[len(prefix):close_idx] body = text[close_idx + len("\n```\n\n"):] ``` Emit `message {reasoning: , text: , …}`. If no prefix is found (reasoning off / no reasoning), emit `message {text: …}` with no `reasoning`. > Robustness: the split is a best-effort parse of a *stable, gateway-owned* > format. If the format ever changes, the fallback is "no reasoning field, full > text" — the app still shows the answer. Verified in M2 against the live > gateway. **App side.** `ReasoningBlock` composable: collapsible, header "💭 Reasoning", monospace body, a **copy** button (matches reference). Rendered **above** the message body. Collapsed by default if long; expanded tap. ## 5.3 Tool output (app controls verbosity) **Requirement:** hermes supports several tool-display modes; the **gateway sends full structured data**, and the **app** chooses how much to show (everything / truncated / nothing). **Gateway side.** Tool activity flows through `tool_progress_callback` → the gateway progress queue → `send_progress_messages` (`gateway/run.py:4603`) → `adapter.send()`. The adapter receives tool lines during a turn. **Adapter → frames.** The adapter classifies tool activity (via turn-state + line format) and emits **structured** frames — not pre-formatted strings: - `tool.start {index, name, preview, args}` — a tool call began. - `tool.progress {index, name, note}` — in-progress update (optional). - `tool.end {index, name, ok, duration, output_preview}` — completed. `index` is a monotonic per-turn counter so the app correlates start→end. `args` is the full argument dict (the app truncates). `output_preview` is a short tail; full tool output is **not** streamed (it lives in agent history and is reachable via search). **App side — the verbosity setting** (Settings → "Tool detail"): - **Everything** — show tool name, full args (collapsible), and output preview. - **Truncated** (default) — show `emoji name: "short preview"` one-liner, collapsible to expand. - **Nothing** — suppress `tool.*` frames entirely (clean chat). Rendering: a `ToolCard` per tool, grouped under the message it belongs to, with a spinner while `tool.end` hasn't arrived, a ✓/✗ on completion, and duration. > Classification detail: the exact way to distinguish a tool-progress `send()` > from a regular `send()`/commentary is confirmed empirically in M2 by running > the real gateway with a test WS client and observing the calls. The adapter > keeps a per-chat turn state machine (turn active, current streaming id, last > tool index) to make the classification deterministic. ## 5.4 Intermediate messages (commentary) **Requirement:** show the agent's interim beats (e.g. "I'll inspect the repo first.") as distinct messages. **Gateway side.** `interim_assistant_callback` → consumer `on_commentary` (`gateway/stream_consumer.py:518`) → delivered as its own message. **Adapter → frames.** `commentary {message_id, text}`. **App side.** Render as a **dimmed / smaller** bubble, visually distinct from final answers (e.g. reduced opacity, no model footer). It reads as a "beat" in the conversation, not a full reply. ## 5.5 Typing indicator `send_typing` → `typing {on:true}`; the app shows an animated indicator in the chat header / above the composer until `typing.stop` or the first `message.start`. ## 5.6 Frame sequence for a typical turn ``` app → message.send {text:"summarize the repo"} srv → typing {on:true} srv → message.start {message_id:m1} srv → message.update {m1, "Let me look…"} (streaming) srv → tool.start {index:1, name:terminal, args:{command:"ls -la"}} srv → tool.end {index:1, name:terminal, ok:true, duration:0.4} srv → commentary {m2, "Found 12 files."} srv → message.start {message_id:m3} srv → message.update {m3, "The repo has…"} srv → message.stop {m3, final_text:"…", reasoning:"…", model:"…", tokens:42} srv → typing {on:false} ```