Fix Hermes adapter integration and recovery

This commit is contained in:
Codex
2026-09-02 13:59:00 -07:00
parent 53ffdb8aa2
commit 5130152335
10 changed files with 900 additions and 357 deletions
+51 -32
View File
@@ -59,7 +59,7 @@ Every error uses one shape with a machine-readable code:
Codes: `unauthorized`, `auth_timeout`, `version_mismatch`, `bad_request`,
`invalid_json`, `unknown_type`, `not_found`, `conversation_not_found`,
`approval_not_found`, `clarify_not_found`, `too_large`, `rate_limited`,
`internal_error`, `not_implemented`.
`run_active`, `internal_error`, `not_implemented`.
Limits: chat text ≤ 64,000 chars; inbound WS frame ≤ 2 MiB (violations get
`too_large`); conversation history fetch ≤ 500 messages.
@@ -98,13 +98,20 @@ Limits: chat text ≤ 64,000 chars; inbound WS frame ≤ 2 MiB (violations get
← { "type": "conversation.history", "conversation_id": "a1b2…", "request_id": "r3",
"messages": [
{ "message_id": "m12", "role": "user", "text": "hey", "ts": "…|null" },
{ "message_id": "m13", "role": "assistant", "text": "hi!", "ts": "…|null" } ] }
{ "message_id": "m13", "role": "assistant", "text": "hi!", "ts": "…|null" } ],
"attachments": [],
"run": null,
"tools": [],
"approvals": [],
"clarifications": [] }
```
History is the authoritative Hermes transcript (`role` is always `user` or
`assistant`). On reconnect, re-open the last-open conversations and resume —
no client-side message cache is needed for correctness. Unknown conversation
→ `conversation_not_found` error.
`assistant`). The other fields form a recoverable snapshot: unexpired
attachments, the active run (if any), latest structured tool states, and
pending approval/clarification requests. On reconnect, replace local state
with this snapshot, then consume new live events. Unknown conversation →
`conversation_not_found` error.
### `conversation.rename`
@@ -168,14 +175,15 @@ Run end:
```json
← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…",
"status": "completed" | "cancelled" | "failed" | "idle",
"status": "completed" | "cancelled" | "failed",
"error": "only on failure" }
```
State machine per assistant turn:
`run.accepted → message.start → (message.delta)* → message.complete → run.finished`.
A turn with no streaming skips `message.start`/`message.delta`. Never infer
state from text — use these events.
A non-streaming turn may skip `message.delta`; `message.start` is always sent
when the run is accepted. Only one run may be active per conversation;
another `chat.send` receives `run_active`. Never infer state from text.
### `run.cancel` — stop an active run
@@ -186,8 +194,8 @@ state from text — use these events.
Cancellation uses Hermes's supported interrupt mechanism (agent interrupt +
run-generation invalidation) — the conversation stays consistent and
resumable. Cancelling with no active run returns `run.finished`
`status:"idle"`.
resumable. A stale `run_id`, missing run, or run that Hermes can no longer
interrupt returns `not_found`; success emits exactly one cancelled event.
---
@@ -200,21 +208,20 @@ Tool activity arrives as `tool.event` messages, completely separate from
```json
{ "type": "tool.event",
"conversation_id": "a1b2…",
"run_id": "8c1f…",
"tool_call_id": "t-1a2b3c4d5e6f",
"tool_name": "web_search",
"status": "running" | "completed" | "failed",
"description": "cats — short preview from the agent (may be null)",
"status": "running" | "completed" | "failed" | "cancelled",
"args_redacted": { "query": "cats" }, // only on "running"; secret-looking keys redacted
"duration_ms": 1234, // only on completion/failure (may be null)
"error": "only on failed, truncated", // only on "failed"
"ts": "..." }
```
Correlate `running` → `completed`/`failed` by `tool_call_id`. Note: the
running event's ID comes from the adapter and the completion event from
Hermes's `post_tool_call` hook; when they differ, correlate by
`(tool_name, conversation)` as a fallback and prefer the completion event's
ID going forward.
Correlate `running` → `completed`/`failed` by `tool_call_id`. Both events use
Hermes's authoritative call ID from the pre/post tool hooks and include the
active `run_id` when one exists. Secret-looking argument keys are redacted
recursively before leaving the server.
No fake "Searching the web…" text is ever injected into `message.*` events.
@@ -228,6 +235,8 @@ When Hermes pauses for a human decision on a dangerous action:
{ "type": "approval.request",
"approval_id": "3d4e5f6070a1",
"session_key": "agent:main:pheby:dm:a1b2…",
"conversation_id": "a1b2…",
"run_id": "8c1f…",
"command": "rm -rf /tmp/build-output",
"description": "Destructive shell command (rm -rf)",
"choices": ["once", "session", "always", "deny"],
@@ -240,11 +249,13 @@ Respond:
→ { "type": "approval.respond", "approval_id": "3d4e5f6070a1",
"choice": "once" | "session" | "always" | "deny",
"reason": "optional free text with deny", "request_id": "r8" }
← { "type": "approval.resolved", "approval_id": "3d4e5f6070a1", "choice": "once", "request_id": "r8" }
// broadcast confirmation (also informs other tabs):
← { "type": "approval.resolved", "approval_id": "3d4e5f6070a1", "choice": "once", "accepted": true }
← { "type": "approval.resolved", "approval_id": "3d4e5f6070a1",
"choice": "once", "accepted": true, "request_id": "r8" }
```
The single resolution event is broadcast to every connected client; the
responding client's `request_id` is included on that broadcast.
Choices map to Hermes semantics: `once` (approve this action), `session`
(approve pattern for this conversation), `always` (also persist), `deny`
(decline; the agent is told NOT to retry). Unknown/stale ID →
@@ -259,6 +270,8 @@ own timeout, so a silently-closed socket can't leave a zombie gate.
{ "type": "clarify.request",
"clarify_id": "c1a2b3d4e5",
"session_key": "agent:main:pheby:dm:a1b2…",
"conversation_id": "a1b2…",
"run_id": "8c1f…",
"question": "Deploy to staging or production?",
"choices": ["staging", "production"], // null ⇒ free text only
"allow_free_text": true,
@@ -269,13 +282,13 @@ Respond (either a choice value or free text):
```json
→ { "type": "clarify.respond", "clarify_id": "c1a2b3d4e5", "response": "production", "request_id": "r9" }
← { "type": "clarify.resolved", "clarify_id": "c1a2b3d4e5", "request_id": "r9" }
← { "type": "clarify.resolved", "clarify_id": "c1a2b3d4e5", "accepted": true } // broadcast
← { "type": "clarify.resolved", "clarify_id": "c1a2b3d4e5",
"accepted": true, "request_id": "r9" }
```
`accepted:false` on the broadcast means Hermes had already resolved/timed out
the prompt. Always render an "Other" affordance — Hermes clarifications
accept free text.
The single resolution event is broadcast. A stale/timed-out ID instead gets
`clarify_not_found`. Always render an "Other" affordance — Hermes
clarifications accept free text.
---
@@ -330,7 +343,7 @@ Authorization: Bearer <PHEBY_SECRET> (or ApiKey <secret>, or X-Pheby-Secre
### List providers + models
```json
→ { "type": "models.list", "request_id": "r10" }
→ { "type": "models.list", "conversation_id": "a1b2…", "request_id": "r10" }
← { "type": "models.snapshot", "request_id": "r10",
"providers": [
{ "slug": "openrouter", "name": "OpenRouter", "is_current": true,
@@ -343,13 +356,17 @@ Authorization: Bearer <PHEBY_SECRET> (or ApiKey <secret>, or X-Pheby-Secre
```
Lists come from Hermes's own credential-aware picker data — nothing is
hardcoded. Models are exactly what the configured providers expose.
hardcoded. Models are exactly what the configured providers expose. The
optional `conversation_id` makes `current_model`, `current_provider`, and
`scope` reflect that conversation's override.
### Read / change current model
```json
→ { "type": "models.current", "request_id": "r11" }
← { "type": "model.current", "request_id": "r11", "model": "z-ai/glm-5.3-flash", "provider": "openrouter", "ts": "..." }
→ { "type": "models.current", "conversation_id": "a1b2…", "request_id": "r11" }
← { "type": "model.current", "request_id": "r11", "model": "z-ai/glm-5.3-flash",
"provider": "openrouter", "scope": "conversation",
"conversation_id": "a1b2…", "ts": "..." }
→ { "type": "model.set", "model": "anthropic/claude-sonnet-4",
"provider": "anthropic", // optional
@@ -369,10 +386,11 @@ conversation the default moved.
## Reasoning effort
```json
→ { "type": "reasoning.current", "request_id": "r13" }
→ { "type": "reasoning.current", "conversation_id": "a1b2…", "request_id": "r13" }
← { "type": "reasoning.snapshot", "request_id": "r13",
"effort": "medium", // current effective effort (may be null = provider default)
"enabled": true, // false ⇒ thinking disabled
"scope": "conversation", "conversation_id": "a1b2…",
"supported_efforts": ["none","minimal","low","medium","high","xhigh","max","ultra"],
"ts": "..." }
@@ -399,8 +417,9 @@ notifications) is pushed as normal `message.*` / `run.*` events even when it
is not a reply to your last request. Reconnection procedure for clients:
1. Reconnect WS, redo `hello`.
2. Re-`conversation.open` the conversations you show; replace local state
with `conversation.history` (authoritative).
2. Re-`conversation.open` the conversations you show; replace local history,
attachments, run/tool state, and pending decisions with its authoritative
snapshot.
3. Re-`models.current` / `reasoning.current` if those views are visible.
4. Live events continue from there.