6be1601ccd
Deliverables now inherit the active assistant draft message ID so clients can render them beside the reply that produced them.
428 lines
16 KiB
Markdown
428 lines
16 KiB
Markdown
# Pheby Protocol v1 — Specification
|
|
|
|
JSON messages over WebSocket, plus one authenticated HTTPS endpoint for
|
|
attachment downloads. Every message (both directions) carries a `"type"`.
|
|
Client→server requests MAY carry a `"request_id"` (any client-chosen string);
|
|
the direct reply echoes it. Server→client events are broadcast to all
|
|
authenticated connections (single-user app — typically one client).
|
|
|
|
Timestamps (`"ts"`) are ISO-8601 UTC. All IDs are opaque strings — the client
|
|
never constructs meaning from them and never sees server filesystem paths.
|
|
|
|
## Handshake
|
|
|
|
Connect to `wss://<host>/ws` (behind Caddy; the plugin itself is plain
|
|
`ws://127.0.0.1:8620/ws`). The **first** client frame must be `hello` within
|
|
10 seconds, or the server closes the socket (`auth_timeout`).
|
|
|
|
### Client → Server `hello`
|
|
|
|
```json
|
|
{
|
|
"type": "hello",
|
|
"secret": "<PHEBY_SECRET>",
|
|
"protocol_version": 1,
|
|
"request_id": "optional"
|
|
}
|
|
```
|
|
|
|
### Server → Client `ready` (success)
|
|
|
|
```json
|
|
{ "type": "ready", "protocol_version": 1, "server": "pheby",
|
|
"ts": "2026-09-02T18:00:00+00:00" }
|
|
```
|
|
|
|
Failure: the server replies with an `error` event (`unauthorized`,
|
|
`auth_timeout`, or `version_mismatch`) and closes. After 5 failed hellos from
|
|
one source address within 60s, further connections are refused (lockout).
|
|
|
|
### Heartbeat
|
|
|
|
```json
|
|
→ { "type": "ping" }
|
|
← { "type": "pong", "ts": "..." }
|
|
```
|
|
|
|
The server also sends WebSocket protocol-level pings (aiohttp `heartbeat=30`).
|
|
|
|
## Errors
|
|
|
|
Every error uses one shape with a machine-readable code:
|
|
|
|
```json
|
|
{ "type": "error", "request_id": "r1",
|
|
"error": { "code": "conversation_not_found", "message": "Conversation not found" },
|
|
"ts": "..." }
|
|
```
|
|
|
|
Codes: `unauthorized`, `auth_timeout`, `version_mismatch`, `bad_request`,
|
|
`invalid_json`, `unknown_type`, `not_found`, `conversation_not_found`,
|
|
`approval_not_found`, `clarify_not_found`, `too_large`, `rate_limited`,
|
|
`run_active`, `internal_error`, `not_implemented`.
|
|
|
|
Limits: chat text ≤ 64,000 chars; inbound WS frame ≤ 2 MiB (violations get
|
|
`too_large`); conversation history fetch ≤ 500 messages.
|
|
|
|
---
|
|
|
|
## Conversations
|
|
|
|
### `conversation.list`
|
|
|
|
```json
|
|
→ { "type": "conversation.list", "request_id": "r1" }
|
|
← { "type": "conversation.snapshot", "request_id": "r1",
|
|
"conversations": [
|
|
{ "conversation_id": "9f1c…", "name": "Project X",
|
|
"session_id": "20260902_101112_ab12cd34", // Hermes session (may be null)
|
|
"last_active": "2026-09-02T17:44:01+00:00", // may be null
|
|
"source": "hermes" } ] }
|
|
```
|
|
|
|
### `conversation.create`
|
|
|
|
```json
|
|
→ { "type": "conversation.create", "name": "New chat", "request_id": "r2" }
|
|
← { "type": "conversation.created", "conversation_id": "a1b2…", "name": "New chat", "request_id": "r2" }
|
|
// plus, broadcast to all clients:
|
|
← { "type": "conversation.updated", "conversation_id": "a1b2…", "name": "New chat" }
|
|
```
|
|
|
|
`name` optional. Conversation IDs are server-generated 32-hex opaque strings.
|
|
|
|
### `conversation.open` — load history (reconnect recovery)
|
|
|
|
```json
|
|
→ { "type": "conversation.open", "conversation_id": "a1b2…", "limit": 200, "request_id": "r3" }
|
|
← { "type": "conversation.history", "conversation_id": "a1b2…", "request_id": "r3",
|
|
"messages": [
|
|
{ "message_id": "m12", "role": "user", "text": "hey", "ts": "…|null" },
|
|
{ "message_id": "m13", "role": "assistant", "text": "hi!", "ts": "…|null" } ],
|
|
"attachments": [],
|
|
"run": null,
|
|
"tools": [],
|
|
"approvals": [],
|
|
"clarifications": [] }
|
|
```
|
|
|
|
History is the authoritative Hermes transcript (`role` is always `user` or
|
|
`assistant`). The other fields form a recoverable snapshot: unexpired
|
|
attachments, the active run (if any), latest structured tool states, and
|
|
pending approval/clarification requests. On reconnect, replace local state
|
|
with this snapshot, then consume new live events. Unknown conversation →
|
|
`conversation_not_found` error.
|
|
|
|
### `conversation.rename`
|
|
|
|
```json
|
|
→ { "type": "conversation.rename", "conversation_id": "a1b2…", "name": "Renamed", "request_id": "r4" }
|
|
← { "type": "conversation.renamed", "conversation_id": "a1b2…", "name": "Renamed", "request_id": "r4" }
|
|
← { "type": "conversation.renamed", "conversation_id": "a1b2…", "name": "Renamed" } // broadcast
|
|
```
|
|
|
|
### `conversation.delete`
|
|
|
|
```json
|
|
→ { "type": "conversation.delete", "conversation_id": "a1b2…", "request_id": "r5" }
|
|
← { "type": "conversation.deleted", "conversation_id": "a1b2…", "request_id": "r5" }
|
|
```
|
|
|
|
Deletes the Hermes session transcript and the routing entry. Attachments
|
|
belonging to the conversation age out on their own 7-day schedule.
|
|
|
|
**Not supported by design:** message editing, per-message deletion,
|
|
regeneration, edit-and-resend. If you need to "undo", send a correction
|
|
message (the agent sees the whole transcript).
|
|
|
|
---
|
|
|
|
## Chat & streaming
|
|
|
|
### `chat.send`
|
|
|
|
```json
|
|
→ { "type": "chat.send", "conversation_id": "a1b2…", "text": "What's the weather?", "request_id": "r6" }
|
|
```
|
|
|
|
### Server → Client run lifecycle
|
|
|
|
```json
|
|
← { "type": "run.accepted", "conversation_id": "a1b2…", "run_id": "8c1f…", "request_id": "r6" }
|
|
← { "type": "message.start", "conversation_id": "a1b2…", "run_id": "8c1f…", "message_id": "draft-8c1f…" }
|
|
```
|
|
|
|
While the agent streams, the server pushes **cumulative** draft text (the
|
|
client can simply replace the bubble's text each time — no delta stitching):
|
|
|
|
```json
|
|
← { "type": "message.delta", "conversation_id": "a1b2…", "message_id": "draft-8c1f…",
|
|
"text": "It's currently 27°C…", "ts": "..." }
|
|
```
|
|
|
|
Completion (final text supersedes the draft — render the final, drop the
|
|
draft):
|
|
|
|
```json
|
|
← { "type": "message.complete", "conversation_id": "a1b2…",
|
|
"message_id": "draft-8c1f…", "text": "…full final answer…", "ts": "..." }
|
|
```
|
|
|
|
`message.complete` with `"kind": "notice"` is a gateway lifecycle/status
|
|
notice rather than conversation content — render or ignore.
|
|
|
|
Run end:
|
|
|
|
```json
|
|
← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…",
|
|
"status": "completed" | "cancelled" | "failed",
|
|
"error": "only on failure" }
|
|
```
|
|
|
|
State machine per assistant turn:
|
|
`run.accepted → message.start → (message.delta)* → message.complete → run.finished`.
|
|
A non-streaming turn may skip `message.delta`; `message.start` is always sent
|
|
when the run is accepted. Only one run may be active per conversation;
|
|
another `chat.send` receives `run_active`. Never infer state from text.
|
|
|
|
### `run.cancel` — stop an active run
|
|
|
|
```json
|
|
→ { "type": "run.cancel", "conversation_id": "a1b2…", "run_id": "8c1f…", "request_id": "r7" }
|
|
← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…", "status": "cancelled" }
|
|
```
|
|
|
|
Cancellation uses Hermes's supported interrupt mechanism (agent interrupt +
|
|
run-generation invalidation) — the conversation stays consistent and
|
|
resumable. A stale `run_id`, missing run, or run that Hermes can no longer
|
|
interrupt returns `not_found`; success emits exactly one cancelled event.
|
|
|
|
---
|
|
|
|
## Tool events (structured — never chat text)
|
|
|
|
Tool activity arrives as `tool.event` messages, completely separate from
|
|
`message.*` chat content. The client renders them as compact tool components
|
|
(ChatGPT-style) attached to the assistant turn.
|
|
|
|
```json
|
|
{ "type": "tool.event",
|
|
"conversation_id": "a1b2…",
|
|
"run_id": "8c1f…",
|
|
"tool_call_id": "t-1a2b3c4d5e6f",
|
|
"tool_name": "web_search",
|
|
"status": "running" | "completed" | "failed" | "cancelled",
|
|
"args_redacted": { "query": "cats" }, // only on "running"; secret-looking keys redacted
|
|
"duration_ms": 1234, // only on completion/failure (may be null)
|
|
"error": "only on failed, truncated", // only on "failed"
|
|
"ts": "..." }
|
|
```
|
|
|
|
Correlate `running` → `completed`/`failed` by `tool_call_id`. Both events use
|
|
Hermes's authoritative call ID from the pre/post tool hooks and include the
|
|
active `run_id` when one exists. Secret-looking argument keys are redacted
|
|
recursively before leaving the server.
|
|
|
|
No fake "Searching the web…" text is ever injected into `message.*` events.
|
|
|
|
---
|
|
|
|
## Approvals
|
|
|
|
When Hermes pauses for a human decision on a dangerous action:
|
|
|
|
```json
|
|
{ "type": "approval.request",
|
|
"approval_id": "3d4e5f6070a1",
|
|
"session_key": "agent:main:pheby:dm:a1b2…",
|
|
"conversation_id": "a1b2…",
|
|
"run_id": "8c1f…",
|
|
"command": "rm -rf /tmp/build-output",
|
|
"description": "Destructive shell command (rm -rf)",
|
|
"choices": ["once", "session", "always", "deny"],
|
|
"ts": "..." }
|
|
```
|
|
|
|
Respond:
|
|
|
|
```json
|
|
→ { "type": "approval.respond", "approval_id": "3d4e5f6070a1",
|
|
"choice": "once" | "session" | "always" | "deny",
|
|
"reason": "optional free text with deny", "request_id": "r8" }
|
|
← { "type": "approval.resolved", "approval_id": "3d4e5f6070a1",
|
|
"choice": "once", "accepted": true, "request_id": "r8" }
|
|
```
|
|
|
|
The single resolution event is broadcast to every connected client; the
|
|
responding client's `request_id` is included on that broadcast.
|
|
|
|
Choices map to Hermes semantics: `once` (approve this action), `session`
|
|
(approve pattern for this conversation), `always` (also persist), `deny`
|
|
(decline; the agent is told NOT to retry). Unknown/stale ID →
|
|
`approval_not_found` error. Hermes itself fails the approval closed after its
|
|
own timeout, so a silently-closed socket can't leave a zombie gate.
|
|
|
|
---
|
|
|
|
## Clarifications / choices
|
|
|
|
```json
|
|
{ "type": "clarify.request",
|
|
"clarify_id": "c1a2b3d4e5",
|
|
"session_key": "agent:main:pheby:dm:a1b2…",
|
|
"conversation_id": "a1b2…",
|
|
"run_id": "8c1f…",
|
|
"question": "Deploy to staging or production?",
|
|
"choices": ["staging", "production"], // null ⇒ free text only
|
|
"allow_free_text": true,
|
|
"ts": "..." }
|
|
```
|
|
|
|
Respond (either a choice value or free text):
|
|
|
|
```json
|
|
→ { "type": "clarify.respond", "clarify_id": "c1a2b3d4e5", "response": "production", "request_id": "r9" }
|
|
← { "type": "clarify.resolved", "clarify_id": "c1a2b3d4e5",
|
|
"accepted": true, "request_id": "r9" }
|
|
```
|
|
|
|
The single resolution event is broadcast. A stale/timed-out ID instead gets
|
|
`clarify_not_found`. Always render an "Other" affordance — Hermes
|
|
clarifications accept free text.
|
|
|
|
---
|
|
|
|
## Attachments (agent → client deliverables)
|
|
|
|
When the agent produces a file (image, document, audio, video… via Hermes's
|
|
normal `MEDIA:` deliverable pipeline), the server copies it into
|
|
adapter-managed storage and broadcasts:
|
|
|
|
```json
|
|
{ "type": "attachment.added",
|
|
"conversation_id": "a1b2…",
|
|
"attachment": {
|
|
"attachment_id": "e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0",
|
|
"filename": "report.pdf",
|
|
"mime_type": "application/pdf",
|
|
"size": 48213,
|
|
"kind": "image" | "voice" | "video" | "audio" | "document",
|
|
"inline_image": false,
|
|
"conversation_id": "a1b2…",
|
|
"message_id": "draft-8c1f…" | null,
|
|
"created_at": "2026-09-02T18:30:00+00:00",
|
|
"expires_at": "2026-09-09T18:30:00+00:00", // null when retention=0
|
|
"download_path": "/attachments/e5f6…" },
|
|
"ts": "..." }
|
|
```
|
|
|
|
Download over **HTTPS** (authenticated — same `PHEBY_SECRET`):
|
|
|
|
```
|
|
GET {download_path}
|
|
Authorization: Bearer <PHEBY_SECRET> (or ApiKey <secret>, or X-Pheby-Secret: <secret>)
|
|
```
|
|
|
|
* `inline_image: true` → `kind == "image"`, safe for an inline preview
|
|
(`BitmapFactory` / `AsyncImage` with the same authenticated GET).
|
|
* Everything else: download and open as an Android document.
|
|
* Unknown / expired / malformed ID → `404` with
|
|
`{"error":{"code":"not_found","message":"Attachment unavailable"}}` —
|
|
no implementation details.
|
|
* **The client can never request arbitrary files** — only registered
|
|
attachment IDs resolve.
|
|
* **Retention:** adapter copies expire after 7 days (configurable) and are
|
|
deleted by an hourly cleanup. Original files the agent produced elsewhere
|
|
on the host are never touched. `expires_at` tells the client when to stop
|
|
offering the download.
|
|
|
|
---
|
|
|
|
## Models
|
|
|
|
### List providers + models
|
|
|
|
```json
|
|
→ { "type": "models.list", "conversation_id": "a1b2…", "request_id": "r10" }
|
|
← { "type": "models.snapshot", "request_id": "r10",
|
|
"providers": [
|
|
{ "slug": "openrouter", "name": "OpenRouter", "is_current": true,
|
|
"models": ["z-ai/glm-5.3-flash", "anthropic/claude-sonnet-4", "…"],
|
|
"total_models": 42 } ],
|
|
"current_model": "z-ai/glm-5.3-flash",
|
|
"current_provider": "openrouter",
|
|
"supported_reasoning_efforts": ["minimal","low","medium","high","xhigh","max","ultra"],
|
|
"ts": "..." }
|
|
```
|
|
|
|
Lists come from Hermes's own credential-aware picker data — nothing is
|
|
hardcoded. Models are exactly what the configured providers expose. The
|
|
optional `conversation_id` makes `current_model`, `current_provider`, and
|
|
`scope` reflect that conversation's override.
|
|
|
|
### Read / change current model
|
|
|
|
```json
|
|
→ { "type": "models.current", "conversation_id": "a1b2…", "request_id": "r11" }
|
|
← { "type": "model.current", "request_id": "r11", "model": "z-ai/glm-5.3-flash",
|
|
"provider": "openrouter", "scope": "conversation",
|
|
"conversation_id": "a1b2…", "ts": "..." }
|
|
|
|
→ { "type": "model.set", "model": "anthropic/claude-sonnet-4",
|
|
"provider": "anthropic", // optional
|
|
"conversation_id": "a1b2…", // present ⇒ session-scoped override
|
|
"request_id": "r12" }
|
|
← { "type": "model.changed", "model": "anthropic/claude-sonnet-4",
|
|
"provider": "anthropic", "scope": "conversation" | "global", "request_id": "r12" }
|
|
// plus broadcast of model.changed (without request_id) to all clients
|
|
```
|
|
|
|
Omit `conversation_id` ⇒ the change is persisted globally (Hermes
|
|
`model.default`). A `model.changed` with `scope:"global"` tells every open
|
|
conversation the default moved.
|
|
|
|
---
|
|
|
|
## Reasoning effort
|
|
|
|
```json
|
|
→ { "type": "reasoning.current", "conversation_id": "a1b2…", "request_id": "r13" }
|
|
← { "type": "reasoning.snapshot", "request_id": "r13",
|
|
"effort": "medium", // current effective effort (may be null = provider default)
|
|
"enabled": true, // false ⇒ thinking disabled
|
|
"scope": "conversation", "conversation_id": "a1b2…",
|
|
"supported_efforts": ["none","minimal","low","medium","high","xhigh","max","ultra"],
|
|
"ts": "..." }
|
|
|
|
→ { "type": "reasoning.set", "effort": "high", "conversation_id": "a1b2…", "request_id": "r14" }
|
|
← { "type": "reasoning.changed", "effort": "high", "scope": "conversation", "request_id": "r14" }
|
|
```
|
|
|
|
`effort: "none"` disables thinking. Invalid values → `bad_request`. Scope
|
|
rules mirror `model.set` (with `conversation_id` ⇒ session override; without ⇒
|
|
global `agent.reasoning_effort`). **Capability note:** Hermes knows *whether*
|
|
a model supports reasoning (models.dev metadata) but does not expose a
|
|
per-provider enum of valid effort values; the listed levels are Hermes's
|
|
canonical set — unsupported levels on a given provider surface as a provider
|
|
error on the next turn, not at set time. This is a documented Hermes
|
|
limitation, not a Pheby guess.
|
|
|
|
---
|
|
|
|
## Unsolicited messages & reconnect behavior
|
|
|
|
The WebSocket stays connected; any Hermes-originated output destined for the
|
|
Pheby platform (scheduled/cron deliveries, background completions,
|
|
notifications) is pushed as normal `message.*` / `run.*` events even when it
|
|
is not a reply to your last request. Reconnection procedure for clients:
|
|
|
|
1. Reconnect WS, redo `hello`.
|
|
2. Re-`conversation.open` the conversations you show; replace local history,
|
|
attachments, run/tool state, and pending decisions with its authoritative
|
|
snapshot.
|
|
3. Re-`models.current` / `reasoning.current` if those views are visible.
|
|
4. Live events continue from there.
|
|
|
|
No external push service exists (no FCM); Android notification behavior is
|
|
the client's responsibility while the socket is down.
|