Files
iris_x_hermes/docs/05-streaming.md
T
ARIA 60ec2b44a7 Auto-threading, history pagination, streaming toggle + tool/reasoning display settings
- Auto-threading (Telegram topic-mode workflow): message.send {auto_thread}
  mints a fresh AI-named thread (instant derived title, LLM upgrade via
  channel.renamed); channel.created {auto:true}; the app jumps into the new
  thread and relocates the optimistic pending bubble.
- history frame: paged full message history for initial channel open /
  scroll-up pagination (reconstructed from the outbox log).
- Streaming on/off: gateway side (display.platforms.android.streaming) plus a
  per-device app toggle (Settings → Streaming); reasoning/model/tokens carried
  on message frames.
- Context menu: long-press (touch) / right-click (desktop) thread affordances
  via a KMP rightClick expect/actual.
- Settings → Reasoning: auto-collapse long reasoning blocks (default on).
- Tool detail: the gateway now always supplies full tool data — it forces
  verbose tool progress (full args → tool.start.args) and captures each
  completed call via the post_tool_call hook (output/duration/ok → tool.end).
  The app reveals the full call + output on expand (Truncated) and
  auto-expands cards in Everything mode.
2026-08-20 16:26:52 +02:00

6.8 KiB

05 — Streaming, Reasoning, Tools, Intermediate Messages

This doc covers the four "agent transparency" features and exactly how each is produced by the gateway and rendered by the app.

The main hermes gateway delivers via the legacy callback path (not the ACP-only event-native render_message_event path). We map the legacy send/edit_message/progress calls to our WS frames.

5.1 Text streaming

Gateway side. The agent's stream_delta_callback feeds a GatewayStreamConsumer (gateway/stream_consumer.py:156). The consumer accumulates text and, at intervals / thresholds, calls:

  • adapter.send(chat_id, text) — first time a bubble is created.
  • adapter.edit_message(chat_id, message_id, text) — subsequent updates (each carries the full accumulated text).

Adapter → frames.

  • First send() of a turn segment → message.start {message_id, role}.
  • Each edit_message() → message.update {message_id, text} (full text).
  • Segment/turn finalization → message.stop {message_id, final_text, reasoning?, model?, tokens?} (or a final message frame when not streaming).

App side. Maintain a live bubble per message_id. On message.update, replace the bubble text (cheap: it's a full snapshot). On message.stop, finalize (attach reasoning/model/tokens footer, stop the cursor). Auto-scroll while the user is at the bottom.

Streaming on/off. Two levels:

  • Gateway side: hermes display.platforms.android.streaming (default follows global). When off, the app just gets one final message frame.
  • App side (per device): Settings → "Streaming" toggle (default on). When off, the app ignores message.start/message.update frames and materializes each reply as a single final message on message.stop (ChatStore.onMessageStop already handles a stop without a live bubble). The gateway keeps streaming — frames are broadcast to all devices, so this is a display preference like tool verbosity, not a wire flag.

5.2 Reasoning (shown before the message)

Requirement: reasoning appears as a block above the answer (like the reference screenshot's "Reasoning:" panel with a copy button).

Gateway side. hermes prepends reasoning to the final response when show_reasoning is enabled (gateway/run.py:20089). The format is stable and chosen by reasoning_style (gateway/display_config.py:37):

  • code (default): 💭 **Reasoning:**\n```\n<reasoning>\n```\n\n<response>
  • blockquote: > 💭 **Reasoning:**\n> …\n\n<response>
  • subtext: -# 💭 Reasoning\n-# …\n\n<response> (Discord-style)

Plugin config. Set for the android platform:

display:
  platforms:
    android:
      show_reasoning: true
      reasoning_style: code     # we split on the code-fence form

Adapter split. In send(), detect the code-style prefix and split:

prefix = "💭 **Reasoning:**\n```\n"
# find the closing "\n```\n\n" after the prefix
reasoning = text[len(prefix):close_idx]
body      = text[close_idx + len("\n```\n\n"):]

Emit message {reasoning: <reasoning>, text: <body>, …}. If no prefix is found (reasoning off / no reasoning), emit message {text: …} with no reasoning.

Robustness: the split is a best-effort parse of a stable, gateway-owned format. If the format ever changes, the fallback is "no reasoning field, full text" — the app still shows the answer. Verified in M2 against the live gateway.

App side. ReasoningBlock composable: collapsible, header "💭 Reasoning", monospace body, a copy button (matches reference). Rendered above the message body. Collapsed by default if long; expanded tap.

5.3 Tool output (app controls verbosity)

Requirement: hermes supports several tool-display modes; the gateway sends full structured data, and the app chooses how much to show (everything / truncated / nothing).

Gateway side. Tool activity flows through tool_progress_callback → the gateway progress queue → send_progress_messages (gateway/run.py:4603) → adapter.send(). The adapter receives tool lines during a turn.

Adapter → frames. The adapter classifies tool activity (via turn-state + line format) and emits structured frames — not pre-formatted strings:

  • tool.start {index, name, preview, args} — a tool call began.
  • tool.progress {index, name, note} — in-progress update (optional).
  • tool.end {index, name, ok, duration, output_preview} — completed.

index is a monotonic per-turn counter so the app correlates start→end. args is the full argument dict (the app truncates). output_preview is a short tail; full tool output is not streamed (it lives in agent history and is reachable via search).

App side — the verbosity setting (Settings → "Tool detail"):

  • Everything — show tool name, full args (collapsible), and output preview.
  • Truncated (default) — show emoji name: "short preview" one-liner, collapsible to expand.
  • Nothing — suppress tool.* frames entirely (clean chat).

Rendering: a ToolCard per tool, grouped under the message it belongs to, with a spinner while tool.end hasn't arrived, a ✓/✗ on completion, and duration.

Classification detail: the exact way to distinguish a tool-progress send() from a regular send()/commentary is confirmed empirically in M2 by running the real gateway with a test WS client and observing the calls. The adapter keeps a per-chat turn state machine (turn active, current streaming id, last tool index) to make the classification deterministic.

5.4 Intermediate messages (commentary)

Requirement: show the agent's interim beats (e.g. "I'll inspect the repo first.") as distinct messages.

Gateway side. interim_assistant_callback → consumer on_commentary (gateway/stream_consumer.py:518) → delivered as its own message.

Adapter → frames. commentary {message_id, text}.

App side. Render as a dimmed / smaller bubble, visually distinct from final answers (e.g. reduced opacity, no model footer). It reads as a "beat" in the conversation, not a full reply.

5.5 Typing indicator

send_typing → typing {on:true}; the app shows an animated indicator in the chat header / above the composer until typing.stop or the first message.start.

5.6 Frame sequence for a typical turn

app  → message.send {text:"summarize the repo"}
srv  → typing {on:true}
srv  → message.start {message_id:m1}
srv  → message.update {m1, "Let me look…"}        (streaming)
srv  → tool.start   {index:1, name:terminal, args:{command:"ls -la"}}
srv  → tool.end     {index:1, name:terminal, ok:true, duration:0.4}
srv  → commentary   {m2, "Found 12 files."}
srv  → message.start {message_id:m3}
srv  → message.update {m3, "The repo has…"}
srv  → message.stop {m3, final_text:"…", reasoning:"…", model:"…", tokens:42}
srv  → typing {on:false}