M0: toolchain, monorepo scaffold, gateway plugin skeleton, CMP app
- gateway-plugin/: android platform plugin (plugin.yaml + adapter.py register(ctx) + no-op AndroidAdapter) + stub modules for M1-M5 - app/: Compose Multiplatform project (shared KMP + androidApp + desktopApp) with Gradle wrapper; builds :androidApp:assembleDebug and :desktopApp:compileKotlin - scripts/guard_hermes_agent.sh + pre-commit hook: fail if hermes-agent/ is staged (read-only reference, never committed) - .gitignore excludes hermes-agent/; docs/ reference library
This commit is contained in:
commit
59acf66c89
49 files changed
+3950
No files matched your search
@@ -0,0 +1,143 @@
|
||||
# 05 — Streaming, Reasoning, Tools, Intermediate Messages
|
||||
|
||||
This doc covers the four "agent transparency" features and exactly how each is
|
||||
produced by the gateway and rendered by the app.
|
||||
|
||||
> The main hermes gateway delivers via the **legacy callback path** (not the
|
||||
> ACP-only event-native `render_message_event` path). We map the legacy
|
||||
> `send`/`edit_message`/progress calls to our WS frames.
|
||||
|
||||
## 5.1 Text streaming
|
||||
|
||||
**Gateway side.** The agent's `stream_delta_callback` feeds a
|
||||
`GatewayStreamConsumer` (`gateway/stream_consumer.py:156`). The consumer
|
||||
accumulates text and, at intervals / thresholds, calls:
|
||||
- `adapter.send(chat_id, text)` — first time a bubble is created.
|
||||
- `adapter.edit_message(chat_id, message_id, text)` — subsequent updates
|
||||
(each carries the **full** accumulated text).
|
||||
|
||||
**Adapter → frames.**
|
||||
- First `send()` of a turn segment → `message.start {message_id, role}`.
|
||||
- Each `edit_message()` → `message.update {message_id, text}` (full text).
|
||||
- Segment/turn finalization → `message.stop {message_id, final_text, reasoning?,
|
||||
model?, tokens?}` (or a final `message` frame when not streaming).
|
||||
|
||||
**App side.** Maintain a live bubble per `message_id`. On `message.update`,
|
||||
replace the bubble text (cheap: it's a full snapshot). On `message.stop`,
|
||||
finalize (attach reasoning/model/tokens footer, stop the cursor). Auto-scroll
|
||||
while the user is at the bottom.
|
||||
|
||||
**Streaming on/off.** Controlled by hermes `display.platforms.android.streaming`
|
||||
(default follows global). When off, the app just gets one final `message` frame.
|
||||
|
||||
## 5.2 Reasoning (shown *before* the message)
|
||||
|
||||
**Requirement:** reasoning appears as a block **above** the answer (like the
|
||||
reference screenshot's "Reasoning:" panel with a copy button).
|
||||
|
||||
**Gateway side.** hermes prepends reasoning to the final response when
|
||||
`show_reasoning` is enabled (`gateway/run.py:20089`). The format is stable and
|
||||
chosen by `reasoning_style` (`gateway/display_config.py:37`):
|
||||
- `code` (default): `💭 **Reasoning:**\n```\n<reasoning>\n```\n\n<response>`
|
||||
- `blockquote`: `> 💭 **Reasoning:**\n> …\n\n<response>`
|
||||
- `subtext`: `-# 💭 Reasoning\n-# …\n\n<response>` (Discord-style)
|
||||
|
||||
**Plugin config.** Set for the `android` platform:
|
||||
```yaml
|
||||
display:
|
||||
platforms:
|
||||
android:
|
||||
show_reasoning: true
|
||||
reasoning_style: code # we split on the code-fence form
|
||||
```
|
||||
|
||||
**Adapter split.** In `send()`, detect the `code`-style prefix and split:
|
||||
```
|
||||
prefix = "💭 **Reasoning:**\n```\n"
|
||||
# find the closing "\n```\n\n" after the prefix
|
||||
reasoning = text[len(prefix):close_idx]
|
||||
body = text[close_idx + len("\n```\n\n"):]
|
||||
```
|
||||
Emit `message {reasoning: <reasoning>, text: <body>, …}`. If no prefix is found
|
||||
(reasoning off / no reasoning), emit `message {text: …}` with no `reasoning`.
|
||||
|
||||
> Robustness: the split is a best-effort parse of a *stable, gateway-owned*
|
||||
> format. If the format ever changes, the fallback is "no reasoning field, full
|
||||
> text" — the app still shows the answer. Verified in M2 against the live
|
||||
> gateway.
|
||||
|
||||
**App side.** `ReasoningBlock` composable: collapsible, header "💭 Reasoning",
|
||||
monospace body, a **copy** button (matches reference). Rendered **above** the
|
||||
message body. Collapsed by default if long; expanded tap.
|
||||
|
||||
## 5.3 Tool output (app controls verbosity)
|
||||
|
||||
**Requirement:** hermes supports several tool-display modes; the **gateway sends
|
||||
full structured data**, and the **app** chooses how much to show
|
||||
(everything / truncated / nothing).
|
||||
|
||||
**Gateway side.** Tool activity flows through `tool_progress_callback` → the
|
||||
gateway progress queue → `send_progress_messages` (`gateway/run.py:4603`) →
|
||||
`adapter.send()`. The adapter receives tool lines during a turn.
|
||||
|
||||
**Adapter → frames.** The adapter classifies tool activity (via turn-state +
|
||||
line format) and emits **structured** frames — not pre-formatted strings:
|
||||
- `tool.start {index, name, preview, args}` — a tool call began.
|
||||
- `tool.progress {index, name, note}` — in-progress update (optional).
|
||||
- `tool.end {index, name, ok, duration, output_preview}` — completed.
|
||||
|
||||
`index` is a monotonic per-turn counter so the app correlates start→end.
|
||||
`args` is the full argument dict (the app truncates). `output_preview` is a
|
||||
short tail; full tool output is **not** streamed (it lives in agent history and
|
||||
is reachable via search).
|
||||
|
||||
**App side — the verbosity setting** (Settings → "Tool detail"):
|
||||
- **Everything** — show tool name, full args (collapsible), and output preview.
|
||||
- **Truncated** (default) — show `emoji name: "short preview"` one-liner,
|
||||
collapsible to expand.
|
||||
- **Nothing** — suppress `tool.*` frames entirely (clean chat).
|
||||
|
||||
Rendering: a `ToolCard` per tool, grouped under the message it belongs to, with
|
||||
a spinner while `tool.end` hasn't arrived, a ✓/✗ on completion, and duration.
|
||||
|
||||
> Classification detail: the exact way to distinguish a tool-progress `send()`
|
||||
> from a regular `send()`/commentary is confirmed empirically in M2 by running
|
||||
> the real gateway with a test WS client and observing the calls. The adapter
|
||||
> keeps a per-chat turn state machine (turn active, current streaming id, last
|
||||
> tool index) to make the classification deterministic.
|
||||
|
||||
## 5.4 Intermediate messages (commentary)
|
||||
|
||||
**Requirement:** show the agent's interim beats (e.g. "I'll inspect the repo
|
||||
first.") as distinct messages.
|
||||
|
||||
**Gateway side.** `interim_assistant_callback` → consumer `on_commentary`
|
||||
(`gateway/stream_consumer.py:518`) → delivered as its own message.
|
||||
|
||||
**Adapter → frames.** `commentary {message_id, text}`.
|
||||
|
||||
**App side.** Render as a **dimmed / smaller** bubble, visually distinct from
|
||||
final answers (e.g. reduced opacity, no model footer). It reads as a "beat" in
|
||||
the conversation, not a full reply.
|
||||
|
||||
## 5.5 Typing indicator
|
||||
|
||||
`send_typing` → `typing {on:true}`; the app shows an animated indicator in the
|
||||
chat header / above the composer until `typing.stop` or the first
|
||||
`message.start`.
|
||||
|
||||
## 5.6 Frame sequence for a typical turn
|
||||
|
||||
```
|
||||
app → message.send {text:"summarize the repo"}
|
||||
srv → typing {on:true}
|
||||
srv → message.start {message_id:m1}
|
||||
srv → message.update {m1, "Let me look…"} (streaming)
|
||||
srv → tool.start {index:1, name:terminal, args:{command:"ls -la"}}
|
||||
srv → tool.end {index:1, name:terminal, ok:true, duration:0.4}
|
||||
srv → commentary {m2, "Found 12 files."}
|
||||
srv → message.start {message_id:m3}
|
||||
srv → message.update {m3, "The repo has…"}
|
||||
srv → message.stop {m3, final_text:"…", reasoning:"…", model:"…", tokens:42}
|
||||
srv → typing {on:false}
|
||||
```
|
||||
Reference in new issue
Block a user