- MarkdownText: streaming branch on the library's StreamingMarkdownState (rememberStreamingMarkdownPipeline) — incremental append, no full re-parse per snapshot; static path keeps retainState=true - client-side reveal: buffer snapshots, reveal at a steady rate (streamCharsPerSecond = 60/smoothness, 300..75 c/s) with rate-relative catch-up (jump when backlog > 2s of reveal time) - keep the streaming renderer alive after message.stop until the reveal catches up (useStreaming gate); artifact card waits for the same - cursor drawn via annotator, never part of the parse input - new 'Streaming speed' setting (0.2-0.8s) in Settings, persisted per device via SecureStore, live-updatable - docs/05-streaming: rate mapping, catch-up, wait-for-reveal semantics
196 lines
9.1 KiB
Markdown
196 lines
9.1 KiB
Markdown
# 05 — Streaming, Reasoning, Tools, Intermediate Messages
|
||
|
||
This doc covers the four "agent transparency" features and exactly how each is
|
||
produced by the gateway and rendered by the app.
|
||
|
||
> The main hermes gateway delivers via the **legacy callback path** (not the
|
||
> ACP-only event-native `render_message_event` path). We map the legacy
|
||
> `send`/`edit_message`/progress calls to our WS frames.
|
||
|
||
## 5.1 Text streaming
|
||
|
||
**Gateway side.** The agent's `stream_delta_callback` feeds a
|
||
`GatewayStreamConsumer` (`gateway/stream_consumer.py:156`). The consumer
|
||
accumulates text and, at intervals / thresholds, calls:
|
||
|
||
- `adapter.send(chat_id, text)` — first time a bubble is created.
|
||
- `adapter.edit_message(chat_id, message_id, text)` — subsequent updates
|
||
(each carries the **full** accumulated text).
|
||
|
||
**Adapter → frames.**
|
||
|
||
- First `send()` of a turn segment → `message.start {message_id, role}`.
|
||
- Each `edit_message()` → `message.update {message_id, text}` (full text).
|
||
- Segment/turn finalization → `message.stop {message_id, final_text, reasoning?,
|
||
model?, tokens?}` (or a final `message` frame when not streaming).
|
||
|
||
**App side.** Maintain a live bubble per `message_id`. On `message.update`,
|
||
replace the bubble text (cheap: it's a full snapshot). On `message.stop`,
|
||
finalize (attach reasoning/model/tokens footer, stop the cursor). Auto-scroll
|
||
while the user is at the bottom.
|
||
|
||
**Smooth streaming (app-side rendering).** Re-parsing + re-laying-out the
|
||
whole bubble on every update made streamed text unreadable (raw-markdown
|
||
flashes, constant reflow). `MarkdownText` therefore renders streaming bubbles
|
||
incrementally:
|
||
|
||
- **Reveal (typewriter):** the latest full snapshot is revealed at a steady
|
||
rate instead of jumping per gateway update. The rate is user-adjustable:
|
||
Settings → "Streaming speed" (`streamSmoothness`, 0.2–0.8 s between visible
|
||
updates, lower = faster; `streamCharsPerSecond` maps it to chars/s:
|
||
`60 / smoothness`, so 0.2 → 300 chars/s, 0.8 → 75 chars/s). Applied on the
|
||
fly; a backlog exceeding ~2 s of reveal time is revealed at once (fast
|
||
model / reconnect catch-up).
|
||
- **Wait for the reveal:** the streaming renderer stays active until the
|
||
reveal catches up, even after `message.stop` — a fast model that dumps the
|
||
whole text in one or two frames still plays out the typewriter instead of
|
||
jumping to the full message. Only then does the bubble switch to the static
|
||
renderer (which also enables the HTML artifact card).
|
||
- **Incremental parse:** the revealed prefix is fed as append chunks
|
||
(`appendChunk`) into the library's `StreamingMarkdownState`
|
||
(`rememberStreamingMarkdownState`, mikepenz 0.44.0) — an append-only parser
|
||
that re-parses only the unstable tail. Settled blocks keep AST identity, so
|
||
Compose never re-lays them out; only the tail re-renders per tick. A
|
||
non-extension snapshot (rewrite) recreates the parser state and re-seeds it.
|
||
- **Prefix-preserving transform:** `preserveNewlinesAsHardBreaksStreaming()`
|
||
is like `preserveNewlinesAsHardBreaks()` but skips the still-incomplete last
|
||
line, so the transform of a prefix is always a prefix of the transform of
|
||
the whole (required for pure-append diffs). The static path keeps the
|
||
original transform plus `retainState = true` (last formatted output stays
|
||
visible during re-parses — no raw flash).
|
||
- **Cursor:** the ▉ is appended to the last text leaf by the annotator (not to
|
||
the parse input, which would break the append diff).
|
||
|
||
The gateway cadence (`edit_interval` / `buffer_threshold`, default 0.8 s /
|
||
24 chars) is unchanged — smoothing happens entirely client-side, so it works
|
||
with any gateway and per device.
|
||
|
||
**Streaming on/off.** Two levels:
|
||
|
||
- **Gateway side:** hermes `display.platforms.iris.streaming` (default
|
||
follows global). When off, the app just gets one final `message` frame.
|
||
- **App side (per device):** Settings → "Streaming" toggle (default on). When
|
||
off, the app ignores `message.start`/`message.update` frames and
|
||
materializes each reply as a single final message on `message.stop`
|
||
(`ChatStore.onMessageStop` already handles a stop without a live bubble).
|
||
The gateway keeps streaming — frames are broadcast to all devices, so this
|
||
is a display preference like tool verbosity, not a wire flag.
|
||
|
||
## 5.2 Reasoning (shown *before* the message)
|
||
|
||
**Requirement:** reasoning appears as a block **above** the answer (like the
|
||
reference screenshot's "Reasoning:" panel with a copy button).
|
||
|
||
**Gateway side.** hermes prepends reasoning to the final response when
|
||
`show_reasoning` is enabled (`gateway/run.py:20089`). The format is stable and
|
||
chosen by `reasoning_style` (`gateway/display_config.py:37`):
|
||
|
||
- `code` (default): `💭 **Reasoning:**\n```\n<reasoning>\n```\n\n<response>`
|
||
- `blockquote`: `> 💭 **Reasoning:**\n> …\n\n<response>`
|
||
- `subtext`: `-# 💭 Reasoning\n-# …\n\n<response>` (Discord-style)
|
||
|
||
**Plugin config.** Set for the `iris` platform:
|
||
|
||
```yaml
|
||
display:
|
||
platforms:
|
||
iris:
|
||
show_reasoning: true
|
||
reasoning_style: code # we split on the code-fence form
|
||
```
|
||
|
||
**Adapter split.** In `send()`, detect the `code`-style prefix and split:
|
||
|
||
```
|
||
prefix = "💭 **Reasoning:**\n```\n"
|
||
# find the closing "\n```\n\n" after the prefix
|
||
reasoning = text[len(prefix):close_idx]
|
||
body = text[close_idx + len("\n```\n\n"):]
|
||
```
|
||
|
||
Emit `message {reasoning: <reasoning>, text: <body>, …}`. If no prefix is found
|
||
(reasoning off / no reasoning), emit `message {text: …}` with no `reasoning`.
|
||
|
||
> Robustness: the split is a best-effort parse of a *stable, gateway-owned*
|
||
> format. If the format ever changes, the fallback is "no reasoning field, full
|
||
> text" — the app still shows the answer. Verified in M2 against the live
|
||
> gateway.
|
||
|
||
**App side.** `ReasoningBlock` composable: collapsible, header "💭 Reasoning",
|
||
monospace body, a **copy** button (matches reference). Rendered **above** the
|
||
message body. Collapsed by default if long; expanded tap.
|
||
|
||
## 5.3 Tool output (app controls verbosity)
|
||
|
||
**Requirement:** hermes supports several tool-display modes; the **gateway sends
|
||
full structured data**, and the **app** chooses how much to show
|
||
(everything / truncated / nothing).
|
||
|
||
**Gateway side.** Tool activity flows through `tool_progress_callback` → the
|
||
gateway progress queue → `send_progress_messages` (`gateway/run.py:4603`) →
|
||
`adapter.send()`. The adapter receives tool lines during a turn.
|
||
|
||
**Adapter → frames.** The adapter classifies tool activity (via turn-state +
|
||
line format) and emits **structured** frames — not pre-formatted strings:
|
||
|
||
- `tool.start {index, name, preview, args}` — a tool call began.
|
||
- `tool.progress {index, name, note}` — in-progress update (optional).
|
||
- `tool.end {index, name, ok, duration, output_preview}` — completed.
|
||
|
||
`index` is a monotonic per-turn counter so the app correlates start→end.
|
||
`args` is the full argument dict (the app truncates). `output_preview` is a
|
||
short tail; full tool output is **not** streamed (it lives in agent history and
|
||
is reachable via search).
|
||
|
||
**App side — the verbosity setting** (Settings → "Tool detail"):
|
||
|
||
- **Everything** — show tool name, full args (collapsible), and output preview.
|
||
- **Truncated** (default) — show `emoji name: "short preview"` one-liner,
|
||
collapsible to expand.
|
||
- **Nothing** — suppress `tool.*` frames entirely (clean chat).
|
||
|
||
Rendering: a `ToolCard` per tool, grouped under the message it belongs to, with
|
||
a spinner while `tool.end` hasn't arrived, a ✓/✗ on completion, and duration.
|
||
|
||
> Classification detail: the exact way to distinguish a tool-progress `send()`
|
||
> from a regular `send()`/commentary is confirmed empirically in M2 by running
|
||
> the real gateway with a test WS client and observing the calls. The adapter
|
||
> keeps a per-chat turn state machine (turn active, current streaming id, last
|
||
> tool index) to make the classification deterministic.
|
||
|
||
## 5.4 Intermediate messages (commentary)
|
||
|
||
**Requirement:** show the agent's interim beats (e.g. "I'll inspect the repo
|
||
first.") as distinct messages.
|
||
|
||
**Gateway side.** `interim_assistant_callback` → consumer `on_commentary`
|
||
(`gateway/stream_consumer.py:518`) → delivered as its own message.
|
||
|
||
**Adapter → frames.** `commentary {message_id, text}`.
|
||
|
||
**App side.** Render as a **dimmed / smaller** bubble, visually distinct from
|
||
final answers (e.g. reduced opacity, no model footer). It reads as a "beat" in
|
||
the conversation, not a full reply.
|
||
|
||
## 5.5 Typing indicator
|
||
|
||
`send_typing` → `typing {on:true}`; the app shows an animated indicator in the
|
||
chat header / above the composer until `typing.stop` or the first
|
||
`message.start`.
|
||
|
||
## 5.6 Frame sequence for a typical turn
|
||
|
||
```
|
||
app → message.send {text:"summarize the repo"}
|
||
srv → typing {on:true}
|
||
srv → message.start {message_id:m1}
|
||
srv → message.update {m1, "Let me look…"} (streaming)
|
||
srv → tool.start {index:1, name:terminal, args:{command:"ls -la"}}
|
||
srv → tool.end {index:1, name:terminal, ok:true, duration:0.4}
|
||
srv → commentary {m2, "Found 12 files."}
|
||
srv → message.start {message_id:m3}
|
||
srv → message.update {m3, "The repo has…"}
|
||
srv → message.stop {m3, final_text:"…", reasoning:"…", model:"…", tokens:42}
|
||
srv → typing {on:false}
|
||
```
|