# Pheby Protocol v1 — Specification JSON messages over WebSocket, plus authenticated HTTPS endpoints for attachment uploads and downloads. Every message (both directions) carries a `"type"`. Client→server requests MAY carry a `"request_id"` (any client-chosen string); the direct reply echoes it. Server→client events are broadcast to all authenticated connections (single-user app — typically one client). Timestamps (`"ts"`) are ISO-8601 UTC. All IDs are opaque strings — the client never constructs meaning from them and never sees server filesystem paths. ## Handshake Connect to `wss:///ws` (behind Caddy; the plugin itself is plain `ws://127.0.0.1:8620/ws`). The **first** client frame must be `hello` within 10 seconds, or the server closes the socket (`auth_timeout`). ### Client → Server `hello` ```json { "type": "hello", "secret": "", "protocol_version": 1, "request_id": "optional" } ``` ### Server → Client `ready` (success) ```json { "type": "ready", "protocol_version": 1, "server": "pheby", "ts": "2026-09-02T18:00:00+00:00" } ``` Failure: the server replies with an `error` event (`unauthorized`, `auth_timeout`, or `version_mismatch`) and closes. After 5 failed hellos from one source address within 60s, further connections are refused (lockout). ### Heartbeat ```json → { "type": "ping" } ← { "type": "pong", "ts": "..." } ``` The server also sends WebSocket protocol-level pings (aiohttp `heartbeat=30`). ## Errors Every error uses one shape with a machine-readable code: ```json { "type": "error", "request_id": "r1", "error": { "code": "conversation_not_found", "message": "Conversation not found" }, "ts": "..." } ``` Codes: `unauthorized`, `auth_timeout`, `version_mismatch`, `bad_request`, `invalid_json`, `unknown_type`, `not_found`, `conversation_not_found`, `approval_not_found`, `clarify_not_found`, `too_large`, `rate_limited`, `run_active`, `internal_error`, `not_implemented`. Limits: chat text ≤ 64,000 chars; inbound WS frame ≤ 2 MiB (violations get `too_large`); conversation history fetch ≤ 500 messages. --- ## Conversations ### `conversation.list` ```json → { "type": "conversation.list", "request_id": "r1" } ← { "type": "conversation.snapshot", "request_id": "r1", "conversations": [ { "conversation_id": "9f1c…", "name": "Project X", "session_id": "20260902_101112_ab12cd34", // Hermes session (may be null) "last_active": "2026-09-02T17:44:01+00:00", // may be null "source": "hermes" } ] } ``` ### `conversation.create` ```json → { "type": "conversation.create", "name": "New chat", "request_id": "r2" } ← { "type": "conversation.created", "conversation_id": "a1b2…", "name": "New chat", "request_id": "r2" } // plus, broadcast to all clients: ← { "type": "conversation.updated", "conversation_id": "a1b2…", "name": "New chat" } ``` `name` optional. Conversation IDs are server-generated 32-hex opaque strings. ### `conversation.open` — load history (reconnect recovery) ```json → { "type": "conversation.open", "conversation_id": "a1b2…", "limit": 200, "request_id": "r3" } ← { "type": "conversation.history", "conversation_id": "a1b2…", "request_id": "r3", "messages": [ { "message_id": "m12", "role": "user", "text": "hey", "ts": "…|null" }, { "message_id": "m13", "role": "assistant", "text": "hi!", "ts": "…|null" } ], "attachments": [], "run": null, "tools": [], "approvals": [], "clarifications": [] } ``` History is the authoritative Hermes transcript (`role` is always `user` or `assistant`). The other fields form a recoverable snapshot: unexpired attachments, the active run (if any), latest structured tool states, and pending approval/clarification requests. On reconnect, replace local state with this snapshot, then consume new live events. Unknown conversation → `conversation_not_found` error. ### `conversation.rename` ```json → { "type": "conversation.rename", "conversation_id": "a1b2…", "name": "Renamed", "request_id": "r4" } ← { "type": "conversation.renamed", "conversation_id": "a1b2…", "name": "Renamed", "request_id": "r4" } ← { "type": "conversation.renamed", "conversation_id": "a1b2…", "name": "Renamed" } // broadcast ``` ### `conversation.delete` ```json → { "type": "conversation.delete", "conversation_id": "a1b2…", "request_id": "r5" } ← { "type": "conversation.deleted", "conversation_id": "a1b2…", "request_id": "r5" } ``` Deletes the Hermes session transcript and the routing entry. Attachments belonging to the conversation age out on their own 7-day schedule. **Not supported by design:** message editing, per-message deletion, regeneration, edit-and-resend. If you need to "undo", send a correction message (the agent sees the whole transcript). --- ## Chat & streaming ### `chat.send` ```json → { "type": "chat.send", "conversation_id": "a1b2…", "text": "What's the weather?", "request_id": "r6" } ``` `attachment_ids` is an optional list of up to 10 distinct inbound upload IDs from `POST /attachments` in the **same conversation**. The adapter rejects unknown, already-sent, or cross-conversation IDs. Upload first, then include all IDs in the single `chat.send`; a rejected active run leaves them retryable. The `text` field is still required and nonblank (for file-only sends, provide a short caption). A successful send anchors the attachments to the user turn and broadcasts `attachment.added`. Small text files are included in agent context; other files remain available to the agent as local media paths. ### Server → Client run lifecycle ```json ← { "type": "run.accepted", "conversation_id": "a1b2…", "run_id": "8c1f…", "request_id": "r6" } ← { "type": "message.start", "conversation_id": "a1b2…", "run_id": "8c1f…", "message_id": "draft-8c1f…" } ``` While the agent streams, the server pushes **cumulative** draft text (the client can simply replace the bubble's text each time — no delta stitching): ```json ← { "type": "message.delta", "conversation_id": "a1b2…", "message_id": "draft-8c1f…", "text": "It's currently 27°C…", "ts": "..." } ``` Completion (final text supersedes the draft — render the final, drop the draft): ```json ← { "type": "message.complete", "conversation_id": "a1b2…", "message_id": "draft-8c1f…", "text": "…full final answer…", "ts": "..." } ``` `message.complete` with `"kind": "notice"` is a gateway lifecycle/status notice rather than conversation content — render or ignore. Run end: ```json ← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…", "status": "completed" | "cancelled" | "failed", "error": "only on failure" } ``` State machine per assistant turn: `run.accepted → message.start → (message.delta)* → message.complete → run.finished`. A non-streaming turn may skip `message.delta`; `message.start` is always sent when the run is accepted. Only one run may be active per conversation; another `chat.send` receives `run_active`. Never infer state from text. ### `run.cancel` — stop an active run ```json → { "type": "run.cancel", "conversation_id": "a1b2…", "run_id": "8c1f…", "request_id": "r7" } ← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…", "status": "cancelled" } ``` Cancellation uses Hermes's supported interrupt mechanism (agent interrupt + run-generation invalidation) — the conversation stays consistent and resumable. A stale `run_id`, missing run, or run that Hermes can no longer interrupt returns `not_found`; success emits exactly one cancelled event. --- ## Tool events (structured — never chat text) Tool activity arrives as `tool.event` messages, completely separate from `message.*` chat content. The client renders them as compact tool components (ChatGPT-style) attached to the assistant turn. ```json { "type": "tool.event", "conversation_id": "a1b2…", "run_id": "8c1f…", "tool_call_id": "t-1a2b3c4d5e6f", "tool_name": "web_search", "status": "running" | "completed" | "failed" | "cancelled", "args_redacted": { "query": "cats" }, // only on "running"; secret-looking keys redacted "duration_ms": 1234, // only on completion/failure (may be null) "error": "only on failed, truncated", // only on "failed" "ts": "..." } ``` Correlate `running` → `completed`/`failed` by `tool_call_id`. Both events use Hermes's authoritative call ID from the pre/post tool hooks and include the active `run_id` when one exists. Secret-looking argument keys are redacted recursively before leaving the server. No fake "Searching the web…" text is ever injected into `message.*` events. --- ## Approvals When Hermes pauses for a human decision on a dangerous action: ```json { "type": "approval.request", "approval_id": "3d4e5f6070a1", "session_key": "agent:main:pheby:dm:a1b2…", "conversation_id": "a1b2…", "run_id": "8c1f…", "command": "rm -rf /tmp/build-output", "description": "Destructive shell command (rm -rf)", "choices": ["once", "session", "always", "deny"], "ts": "..." } ``` Respond: ```json → { "type": "approval.respond", "approval_id": "3d4e5f6070a1", "choice": "once" | "session" | "always" | "deny", "reason": "optional free text with deny", "request_id": "r8" } ← { "type": "approval.resolved", "approval_id": "3d4e5f6070a1", "choice": "once", "accepted": true, "request_id": "r8" } ``` The single resolution event is broadcast to every connected client; the responding client's `request_id` is included on that broadcast. Choices map to Hermes semantics: `once` (approve this action), `session` (approve pattern for this conversation), `always` (also persist), `deny` (decline; the agent is told NOT to retry). Unknown/stale ID → `approval_not_found` error. Hermes itself fails the approval closed after its own timeout, so a silently-closed socket can't leave a zombie gate. --- ## Clarifications / choices ```json { "type": "clarify.request", "clarify_id": "c1a2b3d4e5", "session_key": "agent:main:pheby:dm:a1b2…", "conversation_id": "a1b2…", "run_id": "8c1f…", "question": "Deploy to staging or production?", "choices": ["staging", "production"], // null ⇒ free text only "allow_free_text": true, "ts": "..." } ``` Respond (either a choice value or free text): ```json → { "type": "clarify.respond", "clarify_id": "c1a2b3d4e5", "response": "production", "request_id": "r9" } ← { "type": "clarify.resolved", "clarify_id": "c1a2b3d4e5", "accepted": true, "request_id": "r9" } ``` The single resolution event is broadcast. A stale/timed-out ID instead gets `clarify_not_found`. Always render an "Other" affordance — Hermes clarifications accept free text. --- ## Auto-approve (YOLO) YOLO is always conversation-scoped. It skips ordinary dangerous-command approval prompts for that Hermes session, while Hermes hard deny rules and hardline safety floors still apply. ```json → { "type": "yolo.current", "conversation_id": "a1b2…", "request_id": "r15" } ← { "type": "yolo.snapshot", "enabled": false, "scope": "conversation", "conversation_id": "a1b2…", "request_id": "r15" } → { "type": "yolo.set", "conversation_id": "a1b2…", "enabled": true, "request_id": "r16" } ← { "type": "yolo.changed", "enabled": true, "scope": "conversation", "conversation_id": "a1b2…", "request_id": "r16" } // plus broadcast of yolo.changed (without request_id) to all clients ``` `enabled` must be a JSON boolean. Unknown conversations return `conversation_not_found`. The adapter also persists the flag in Hermes session metadata so it can be restored after the gateway restarts. --- ## Attachments (both directions) ### Client → agent uploads ``` POST /attachments?conversation_id=a1b2…&filename=notes.md Authorization: Bearer Content-Type: text/markdown # Raw file bytes, not multipart or JSON ``` Returns `201 {"attachment": {"attachment_id": "<32 hex chars>", ...}}`. The descriptor contains `direction: "inbound"`, `message_id: null`, and the same fields as agent deliverables below. An upload is **not** broadcast or listed in conversation history until `chat.send` successfully claims it; an unused upload expires with normal retention. Limit: 64 MiB per file. Missing credentials → 401; bad conversation ID or empty body → 400; oversize body → 413. The file picker may select up to 10 files per message. The server keeps only registered copies in adapter-owned storage and never trusts a client-provided path. Markdown and other small `text/*` files (up to 100 KiB) are inlined into the agent's turn; the conversation history returns just the user's original message text. ### Agent → client deliverables When the agent produces a file (image, document, audio, video… via Hermes's normal `MEDIA:` deliverable pipeline), the server copies it into adapter-managed storage and broadcasts: ```json { "type": "attachment.added", "conversation_id": "a1b2…", "attachment": { "attachment_id": "e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0", "filename": "report.pdf", "mime_type": "application/pdf", "size": 48213, "kind": "image" | "voice" | "video" | "audio" | "document", "inline_image": false, "direction": "outbound", "conversation_id": "a1b2…", "message_id": "draft-8c1f…" | null, "created_at": "2026-09-02T18:30:00+00:00", "expires_at": "2026-09-09T18:30:00+00:00", // null when retention=0 "download_path": "/attachments/e5f6…" }, "ts": "..." } ``` Download over **HTTPS** (authenticated — same `PHEBY_SECRET`): ``` GET {download_path} Authorization: Bearer (or ApiKey , or X-Pheby-Secret: ) ``` * `inline_image: true` → `kind == "image"`, safe for an inline preview (`BitmapFactory` / `AsyncImage` with the same authenticated GET). * Everything else: download and open as an Android document. * Unknown / expired / malformed ID → `404` with `{"error":{"code":"not_found","message":"Attachment unavailable"}}` — no implementation details. * **The client can never request arbitrary files** — only registered attachment IDs resolve. * **Retention:** adapter copies expire after 7 days (configurable) and are deleted by an hourly cleanup. Original files the agent produced elsewhere on the host are never touched. `expires_at` tells the client when to stop offering the download. --- ## Models ### List providers + models ```json → { "type": "models.list", "conversation_id": "a1b2…", "request_id": "r10" } ← { "type": "models.snapshot", "request_id": "r10", "providers": [ { "slug": "openrouter", "name": "OpenRouter", "is_current": true, "models": ["z-ai/glm-5.3-flash", "anthropic/claude-sonnet-4", "…"], "total_models": 42 } ], "current_model": "z-ai/glm-5.3-flash", "current_provider": "openrouter", "supported_reasoning_efforts": ["minimal","low","medium","high","xhigh","max","ultra"], "ts": "..." } ``` Lists come from Hermes's own credential-aware picker data — nothing is hardcoded. Models are exactly what the configured providers expose. The optional `conversation_id` makes `current_model`, `current_provider`, and `scope` reflect that conversation's override. ### Read / change current model ```json → { "type": "models.current", "conversation_id": "a1b2…", "request_id": "r11" } ← { "type": "model.current", "request_id": "r11", "model": "z-ai/glm-5.3-flash", "provider": "openrouter", "scope": "conversation", "conversation_id": "a1b2…", "ts": "..." } → { "type": "model.set", "model": "anthropic/claude-sonnet-4", "provider": "anthropic", // optional "conversation_id": "a1b2…", // present ⇒ session-scoped override "request_id": "r12" } ← { "type": "model.changed", "model": "anthropic/claude-sonnet-4", "provider": "anthropic", "scope": "conversation" | "global", "request_id": "r12" } // plus broadcast of model.changed (without request_id) to all clients ``` Omit `conversation_id` ⇒ the change is persisted globally (Hermes `model.default`). A `model.changed` with `scope:"global"` tells every open conversation the default moved. --- ## Reasoning effort ```json → { "type": "reasoning.current", "conversation_id": "a1b2…", "request_id": "r13" } ← { "type": "reasoning.snapshot", "request_id": "r13", "effort": "medium", // current effective effort (may be null = provider default) "enabled": true, // false ⇒ thinking disabled "scope": "conversation", "conversation_id": "a1b2…", "supported_efforts": ["none","minimal","low","medium","high","xhigh","max","ultra"], "ts": "..." } → { "type": "reasoning.set", "effort": "high", "conversation_id": "a1b2…", "request_id": "r14" } ← { "type": "reasoning.changed", "effort": "high", "scope": "conversation", "request_id": "r14" } ``` `effort: "none"` disables thinking. Invalid values → `bad_request`. Scope rules mirror `model.set` (with `conversation_id` ⇒ session override; without ⇒ global `agent.reasoning_effort`). **Capability note:** Hermes knows *whether* a model supports reasoning (models.dev metadata) but does not expose a per-provider enum of valid effort values; the listed levels are Hermes's canonical set — unsupported levels on a given provider surface as a provider error on the next turn, not at set time. This is a documented Hermes limitation, not a Pheby guess. --- ## Unsolicited messages & reconnect behavior The WebSocket stays connected; any Hermes-originated output destined for the Pheby platform (scheduled/cron deliveries, background completions, notifications) is pushed as normal `message.*` / `run.*` events even when it is not a reply to your last request. Reconnection procedure for clients: 1. Reconnect WS, redo `hello`. 2. Re-`conversation.open` the conversations you show; replace local history, attachments, run/tool state, and pending decisions with its authoritative snapshot. 3. Re-`models.current` / `reasoning.current` if those views are visible. 4. Live events continue from there. No external push service exists (no FCM); Android notification behavior is the client's responsibility while the socket is down.