M0: toolchain, monorepo scaffold, gateway plugin skeleton, CMP app

- gateway-plugin/: android platform plugin (plugin.yaml + adapter.py
  register(ctx) + no-op AndroidAdapter) + stub modules for M1-M5
- app/: Compose Multiplatform project (shared KMP + androidApp +
  desktopApp) with Gradle wrapper; builds :androidApp:assembleDebug
  and :desktopApp:compileKotlin
- scripts/guard_hermes_agent.sh + pre-commit hook: fail if hermes-agent/
  is staged (read-only reference, never committed)
- .gitignore excludes hermes-agent/; docs/ reference library
This commit is contained in:
ARIA committed 2026-08-19 11:27:02 +02:00
commit 59acf66c89
49 files changed
+3950

No files matched your search

+143
View File
@@ -0,0 +1,143 @@
# 05 — Streaming, Reasoning, Tools, Intermediate Messages
This doc covers the four "agent transparency" features and exactly how each is
produced by the gateway and rendered by the app.
> The main hermes gateway delivers via the **legacy callback path** (not the
> ACP-only event-native `render_message_event` path). We map the legacy
> `send`/`edit_message`/progress calls to our WS frames.
## 5.1 Text streaming
**Gateway side.** The agent's `stream_delta_callback` feeds a
`GatewayStreamConsumer` (`gateway/stream_consumer.py:156`). The consumer
accumulates text and, at intervals / thresholds, calls:
- `adapter.send(chat_id, text)` — first time a bubble is created.
- `adapter.edit_message(chat_id, message_id, text)` — subsequent updates
(each carries the **full** accumulated text).
**Adapter → frames.**
- First `send()` of a turn segment → `message.start {message_id, role}`.
- Each `edit_message()` → `message.update {message_id, text}` (full text).
- Segment/turn finalization → `message.stop {message_id, final_text, reasoning?,
model?, tokens?}` (or a final `message` frame when not streaming).
**App side.** Maintain a live bubble per `message_id`. On `message.update`,
replace the bubble text (cheap: it's a full snapshot). On `message.stop`,
finalize (attach reasoning/model/tokens footer, stop the cursor). Auto-scroll
while the user is at the bottom.
**Streaming on/off.** Controlled by hermes `display.platforms.android.streaming`
(default follows global). When off, the app just gets one final `message` frame.
## 5.2 Reasoning (shown *before* the message)
**Requirement:** reasoning appears as a block **above** the answer (like the
reference screenshot's "Reasoning:" panel with a copy button).
**Gateway side.** hermes prepends reasoning to the final response when
`show_reasoning` is enabled (`gateway/run.py:20089`). The format is stable and
chosen by `reasoning_style` (`gateway/display_config.py:37`):
- `code` (default): `💭 **Reasoning:**\n```\n<reasoning>\n```\n\n<response>`
- `blockquote`: `> 💭 **Reasoning:**\n> …\n\n<response>`
- `subtext`: `-# 💭 Reasoning\n-# …\n\n<response>` (Discord-style)
**Plugin config.** Set for the `android` platform:
```yaml
display:
platforms:
android:
show_reasoning: true
reasoning_style: code # we split on the code-fence form
```
**Adapter split.** In `send()`, detect the `code`-style prefix and split:
```
prefix = "💭 **Reasoning:**\n```\n"
# find the closing "\n```\n\n" after the prefix
reasoning = text[len(prefix):close_idx]
body = text[close_idx + len("\n```\n\n"):]
```
Emit `message {reasoning: <reasoning>, text: <body>, …}`. If no prefix is found
(reasoning off / no reasoning), emit `message {text: …}` with no `reasoning`.
> Robustness: the split is a best-effort parse of a *stable, gateway-owned*
> format. If the format ever changes, the fallback is "no reasoning field, full
> text" — the app still shows the answer. Verified in M2 against the live
> gateway.
**App side.** `ReasoningBlock` composable: collapsible, header "💭 Reasoning",
monospace body, a **copy** button (matches reference). Rendered **above** the
message body. Collapsed by default if long; expanded tap.
## 5.3 Tool output (app controls verbosity)
**Requirement:** hermes supports several tool-display modes; the **gateway sends
full structured data**, and the **app** chooses how much to show
(everything / truncated / nothing).
**Gateway side.** Tool activity flows through `tool_progress_callback` → the
gateway progress queue → `send_progress_messages` (`gateway/run.py:4603`) →
`adapter.send()`. The adapter receives tool lines during a turn.
**Adapter → frames.** The adapter classifies tool activity (via turn-state +
line format) and emits **structured** frames — not pre-formatted strings:
- `tool.start {index, name, preview, args}` — a tool call began.
- `tool.progress {index, name, note}` — in-progress update (optional).
- `tool.end {index, name, ok, duration, output_preview}` — completed.
`index` is a monotonic per-turn counter so the app correlates start→end.
`args` is the full argument dict (the app truncates). `output_preview` is a
short tail; full tool output is **not** streamed (it lives in agent history and
is reachable via search).
**App side — the verbosity setting** (Settings → "Tool detail"):
- **Everything** — show tool name, full args (collapsible), and output preview.
- **Truncated** (default) — show `emoji name: "short preview"` one-liner,
collapsible to expand.
- **Nothing** — suppress `tool.*` frames entirely (clean chat).
Rendering: a `ToolCard` per tool, grouped under the message it belongs to, with
a spinner while `tool.end` hasn't arrived, a ✓/✗ on completion, and duration.
> Classification detail: the exact way to distinguish a tool-progress `send()`
> from a regular `send()`/commentary is confirmed empirically in M2 by running
> the real gateway with a test WS client and observing the calls. The adapter
> keeps a per-chat turn state machine (turn active, current streaming id, last
> tool index) to make the classification deterministic.
## 5.4 Intermediate messages (commentary)
**Requirement:** show the agent's interim beats (e.g. "I'll inspect the repo
first.") as distinct messages.
**Gateway side.** `interim_assistant_callback` → consumer `on_commentary`
(`gateway/stream_consumer.py:518`) → delivered as its own message.
**Adapter → frames.** `commentary {message_id, text}`.
**App side.** Render as a **dimmed / smaller** bubble, visually distinct from
final answers (e.g. reduced opacity, no model footer). It reads as a "beat" in
the conversation, not a full reply.
## 5.5 Typing indicator
`send_typing` → `typing {on:true}`; the app shows an animated indicator in the
chat header / above the composer until `typing.stop` or the first
`message.start`.
## 5.6 Frame sequence for a typical turn
```
app → message.send {text:"summarize the repo"}
srv → typing {on:true}
srv → message.start {message_id:m1}
srv → message.update {m1, "Let me look…"} (streaming)
srv → tool.start {index:1, name:terminal, args:{command:"ls -la"}}
srv → tool.end {index:1, name:terminal, ok:true, duration:0.4}
srv → commentary {m2, "Found 12 files."}
srv → message.start {message_id:m3}
srv → message.update {m3, "The repo has…"}
srv → message.stop {m3, final_text:"…", reasoning:"…", model:"…", tokens:42}
srv → typing {on:false}
```