# Pheby Protocol v1 — Specification JSON messages over WebSocket, plus one authenticated HTTPS endpoint for attachment downloads. Every message (both directions) carries a `"type"`. Client→server requests MAY carry a `"request_id"` (any client-chosen string); the direct reply echoes it. Server→client events are broadcast to all authenticated connections (single-user app — typically one client). Timestamps (`"ts"`) are ISO-8601 UTC. All IDs are opaque strings — the client never constructs meaning from them and never sees server filesystem paths. ## Handshake Connect to `wss:///ws` (behind Caddy; the plugin itself is plain `ws://127.0.0.1:8620/ws`). The **first** client frame must be `hello` within 10 seconds, or the server closes the socket (`auth_timeout`). ### Client → Server `hello` ```json { "type": "hello", "secret": "", "protocol_version": 1, "request_id": "optional" } ``` ### Server → Client `ready` (success) ```json { "type": "ready", "protocol_version": 1, "server": "pheby", "ts": "2026-09-02T18:00:00+00:00" } ``` Failure: the server replies with an `error` event (`unauthorized`, `auth_timeout`, or `version_mismatch`) and closes. After 5 failed hellos from one source address within 60s, further connections are refused (lockout). ### Heartbeat ```json → { "type": "ping" } ← { "type": "pong", "ts": "..." } ``` The server also sends WebSocket protocol-level pings (aiohttp `heartbeat=30`). ## Errors Every error uses one shape with a machine-readable code: ```json { "type": "error", "request_id": "r1", "error": { "code": "conversation_not_found", "message": "Conversation not found" }, "ts": "..." } ``` Codes: `unauthorized`, `auth_timeout`, `version_mismatch`, `bad_request`, `invalid_json`, `unknown_type`, `not_found`, `conversation_not_found`, `approval_not_found`, `clarify_not_found`, `too_large`, `rate_limited`, `internal_error`, `not_implemented`. Limits: chat text ≤ 64,000 chars; inbound WS frame ≤ 2 MiB (violations get `too_large`); conversation history fetch ≤ 500 messages. --- ## Conversations ### `conversation.list` ```json → { "type": "conversation.list", "request_id": "r1" } ← { "type": "conversation.snapshot", "request_id": "r1", "conversations": [ { "conversation_id": "9f1c…", "name": "Project X", "session_id": "20260902_101112_ab12cd34", // Hermes session (may be null) "last_active": "2026-09-02T17:44:01+00:00", // may be null "source": "hermes" } ] } ``` ### `conversation.create` ```json → { "type": "conversation.create", "name": "New chat", "request_id": "r2" } ← { "type": "conversation.created", "conversation_id": "a1b2…", "name": "New chat", "request_id": "r2" } // plus, broadcast to all clients: ← { "type": "conversation.updated", "conversation_id": "a1b2…", "name": "New chat" } ``` `name` optional. Conversation IDs are server-generated 32-hex opaque strings. ### `conversation.open` — load history (reconnect recovery) ```json → { "type": "conversation.open", "conversation_id": "a1b2…", "limit": 200, "request_id": "r3" } ← { "type": "conversation.history", "conversation_id": "a1b2…", "request_id": "r3", "messages": [ { "message_id": "m12", "role": "user", "text": "hey", "ts": "…|null" }, { "message_id": "m13", "role": "assistant", "text": "hi!", "ts": "…|null" } ] } ``` History is the authoritative Hermes transcript (`role` is always `user` or `assistant`). On reconnect, re-open the last-open conversations and resume — no client-side message cache is needed for correctness. Unknown conversation → `conversation_not_found` error. ### `conversation.rename` ```json → { "type": "conversation.rename", "conversation_id": "a1b2…", "name": "Renamed", "request_id": "r4" } ← { "type": "conversation.renamed", "conversation_id": "a1b2…", "name": "Renamed", "request_id": "r4" } ← { "type": "conversation.renamed", "conversation_id": "a1b2…", "name": "Renamed" } // broadcast ``` ### `conversation.delete` ```json → { "type": "conversation.delete", "conversation_id": "a1b2…", "request_id": "r5" } ← { "type": "conversation.deleted", "conversation_id": "a1b2…", "request_id": "r5" } ``` Deletes the Hermes session transcript and the routing entry. Attachments belonging to the conversation age out on their own 7-day schedule. **Not supported by design:** message editing, per-message deletion, regeneration, edit-and-resend. If you need to "undo", send a correction message (the agent sees the whole transcript). --- ## Chat & streaming ### `chat.send` ```json → { "type": "chat.send", "conversation_id": "a1b2…", "text": "What's the weather?", "request_id": "r6" } ``` ### Server → Client run lifecycle ```json ← { "type": "run.accepted", "conversation_id": "a1b2…", "run_id": "8c1f…", "request_id": "r6" } ← { "type": "message.start", "conversation_id": "a1b2…", "run_id": "8c1f…", "message_id": "draft-8c1f…" } ``` While the agent streams, the server pushes **cumulative** draft text (the client can simply replace the bubble's text each time — no delta stitching): ```json ← { "type": "message.delta", "conversation_id": "a1b2…", "message_id": "draft-8c1f…", "text": "It's currently 27°C…", "ts": "..." } ``` Completion (final text supersedes the draft — render the final, drop the draft): ```json ← { "type": "message.complete", "conversation_id": "a1b2…", "message_id": "draft-8c1f…", "text": "…full final answer…", "ts": "..." } ``` `message.complete` with `"kind": "notice"` is a gateway lifecycle/status notice rather than conversation content — render or ignore. Run end: ```json ← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…", "status": "completed" | "cancelled" | "failed" | "idle", "error": "only on failure" } ``` State machine per assistant turn: `run.accepted → message.start → (message.delta)* → message.complete → run.finished`. A turn with no streaming skips `message.start`/`message.delta`. Never infer state from text — use these events. ### `run.cancel` — stop an active run ```json → { "type": "run.cancel", "conversation_id": "a1b2…", "run_id": "8c1f…", "request_id": "r7" } ← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…", "status": "cancelled" } ``` Cancellation uses Hermes's supported interrupt mechanism (agent interrupt + run-generation invalidation) — the conversation stays consistent and resumable. Cancelling with no active run returns `run.finished` `status:"idle"`. --- ## Tool events (structured — never chat text) Tool activity arrives as `tool.event` messages, completely separate from `message.*` chat content. The client renders them as compact tool components (ChatGPT-style) attached to the assistant turn. ```json { "type": "tool.event", "conversation_id": "a1b2…", "tool_call_id": "t-1a2b3c4d5e6f", "tool_name": "web_search", "status": "running" | "completed" | "failed", "description": "cats — short preview from the agent (may be null)", "args_redacted": { "query": "cats" }, // only on "running"; secret-looking keys redacted "duration_ms": 1234, // only on completion/failure (may be null) "error": "only on failed, truncated", // only on "failed" "ts": "..." } ``` Correlate `running` → `completed`/`failed` by `tool_call_id`. Note: the running event's ID comes from the adapter and the completion event from Hermes's `post_tool_call` hook; when they differ, correlate by `(tool_name, conversation)` as a fallback and prefer the completion event's ID going forward. No fake "Searching the web…" text is ever injected into `message.*` events. --- ## Approvals When Hermes pauses for a human decision on a dangerous action: ```json { "type": "approval.request", "approval_id": "3d4e5f6070a1", "session_key": "agent:main:pheby:dm:a1b2…", "command": "rm -rf /tmp/build-output", "description": "Destructive shell command (rm -rf)", "choices": ["once", "session", "always", "deny"], "ts": "..." } ``` Respond: ```json → { "type": "approval.respond", "approval_id": "3d4e5f6070a1", "choice": "once" | "session" | "always" | "deny", "reason": "optional free text with deny", "request_id": "r8" } ← { "type": "approval.resolved", "approval_id": "3d4e5f6070a1", "choice": "once", "request_id": "r8" } // broadcast confirmation (also informs other tabs): ← { "type": "approval.resolved", "approval_id": "3d4e5f6070a1", "choice": "once", "accepted": true } ``` Choices map to Hermes semantics: `once` (approve this action), `session` (approve pattern for this conversation), `always` (also persist), `deny` (decline; the agent is told NOT to retry). Unknown/stale ID → `approval_not_found` error. Hermes itself fails the approval closed after its own timeout, so a silently-closed socket can't leave a zombie gate. --- ## Clarifications / choices ```json { "type": "clarify.request", "clarify_id": "c1a2b3d4e5", "session_key": "agent:main:pheby:dm:a1b2…", "question": "Deploy to staging or production?", "choices": ["staging", "production"], // null ⇒ free text only "allow_free_text": true, "ts": "..." } ``` Respond (either a choice value or free text): ```json → { "type": "clarify.respond", "clarify_id": "c1a2b3d4e5", "response": "production", "request_id": "r9" } ← { "type": "clarify.resolved", "clarify_id": "c1a2b3d4e5", "request_id": "r9" } ← { "type": "clarify.resolved", "clarify_id": "c1a2b3d4e5", "accepted": true } // broadcast ``` `accepted:false` on the broadcast means Hermes had already resolved/timed out the prompt. Always render an "Other" affordance — Hermes clarifications accept free text. --- ## Attachments (agent → client deliverables) When the agent produces a file (image, document, audio, video… via Hermes's normal `MEDIA:` deliverable pipeline), the server copies it into adapter-managed storage and broadcasts: ```json { "type": "attachment.added", "conversation_id": "a1b2…", "attachment": { "attachment_id": "e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0", "filename": "report.pdf", "mime_type": "application/pdf", "size": 48213, "kind": "image" | "voice" | "video" | "audio" | "document", "inline_image": false, "conversation_id": "a1b2…", "message_id": null, "created_at": "2026-09-02T18:30:00+00:00", "expires_at": "2026-09-09T18:30:00+00:00", // null when retention=0 "download_path": "/attachments/e5f6…" }, "ts": "..." } ``` Download over **HTTPS** (authenticated — same `PHEBY_SECRET`): ``` GET {download_path} Authorization: Bearer (or ApiKey , or X-Pheby-Secret: ) ``` * `inline_image: true` → `kind == "image"`, safe for an inline preview (`BitmapFactory` / `AsyncImage` with the same authenticated GET). * Everything else: download and open as an Android document. * Unknown / expired / malformed ID → `404` with `{"error":{"code":"not_found","message":"Attachment unavailable"}}` — no implementation details. * **The client can never request arbitrary files** — only registered attachment IDs resolve. * **Retention:** adapter copies expire after 7 days (configurable) and are deleted by an hourly cleanup. Original files the agent produced elsewhere on the host are never touched. `expires_at` tells the client when to stop offering the download. --- ## Models ### List providers + models ```json → { "type": "models.list", "request_id": "r10" } ← { "type": "models.snapshot", "request_id": "r10", "providers": [ { "slug": "openrouter", "name": "OpenRouter", "is_current": true, "models": ["z-ai/glm-5.3-flash", "anthropic/claude-sonnet-4", "…"], "total_models": 42 } ], "current_model": "z-ai/glm-5.3-flash", "current_provider": "openrouter", "supported_reasoning_efforts": ["minimal","low","medium","high","xhigh","max","ultra"], "ts": "..." } ``` Lists come from Hermes's own credential-aware picker data — nothing is hardcoded. Models are exactly what the configured providers expose. ### Read / change current model ```json → { "type": "models.current", "request_id": "r11" } ← { "type": "model.current", "request_id": "r11", "model": "z-ai/glm-5.3-flash", "provider": "openrouter", "ts": "..." } → { "type": "model.set", "model": "anthropic/claude-sonnet-4", "provider": "anthropic", // optional "conversation_id": "a1b2…", // present ⇒ session-scoped override "request_id": "r12" } ← { "type": "model.changed", "model": "anthropic/claude-sonnet-4", "provider": "anthropic", "scope": "conversation" | "global", "request_id": "r12" } // plus broadcast of model.changed (without request_id) to all clients ``` Omit `conversation_id` ⇒ the change is persisted globally (Hermes `model.default`). A `model.changed` with `scope:"global"` tells every open conversation the default moved. --- ## Reasoning effort ```json → { "type": "reasoning.current", "request_id": "r13" } ← { "type": "reasoning.snapshot", "request_id": "r13", "effort": "medium", // current effective effort (may be null = provider default) "enabled": true, // false ⇒ thinking disabled "supported_efforts": ["none","minimal","low","medium","high","xhigh","max","ultra"], "ts": "..." } → { "type": "reasoning.set", "effort": "high", "conversation_id": "a1b2…", "request_id": "r14" } ← { "type": "reasoning.changed", "effort": "high", "scope": "conversation", "request_id": "r14" } ``` `effort: "none"` disables thinking. Invalid values → `bad_request`. Scope rules mirror `model.set` (with `conversation_id` ⇒ session override; without ⇒ global `agent.reasoning_effort`). **Capability note:** Hermes knows *whether* a model supports reasoning (models.dev metadata) but does not expose a per-provider enum of valid effort values; the listed levels are Hermes's canonical set — unsupported levels on a given provider surface as a provider error on the next turn, not at set time. This is a documented Hermes limitation, not a Pheby guess. --- ## Unsolicited messages & reconnect behavior The WebSocket stays connected; any Hermes-originated output destined for the Pheby platform (scheduled/cron deliveries, background completions, notifications) is pushed as normal `message.*` / `run.*` events even when it is not a reply to your last request. Reconnection procedure for clients: 1. Reconnect WS, redo `hello`. 2. Re-`conversation.open` the conversations you show; replace local state with `conversation.history` (authoritative). 3. Re-`models.current` / `reasoning.current` if those views are visible. 4. Live events continue from there. No external push service exists (no FCM); Android notification behavior is the client's responsibility while the socket is down.