18 KiB
Pheby Protocol v1 — Specification
JSON messages over WebSocket, plus authenticated HTTPS endpoints for
attachment uploads and downloads. Every message (both directions) carries a "type".
Client→server requests MAY carry a "request_id" (any client-chosen string);
the direct reply echoes it. Server→client events are broadcast to all
authenticated connections (single-user app — typically one client).
Timestamps ("ts") are ISO-8601 UTC. All IDs are opaque strings — the client
never constructs meaning from them and never sees server filesystem paths.
Handshake
Connect to wss://<host>/ws (behind Caddy; the plugin itself is plain
ws://127.0.0.1:8620/ws). The first client frame must be hello within
10 seconds, or the server closes the socket (auth_timeout).
Client → Server hello
{
"type": "hello",
"secret": "<PHEBY_SECRET>",
"protocol_version": 1,
"request_id": "optional"
}
Server → Client ready (success)
{ "type": "ready", "protocol_version": 1, "server": "pheby",
"ts": "2026-09-02T18:00:00+00:00" }
Failure: the server replies with an error event (unauthorized,
auth_timeout, or version_mismatch) and closes. After 5 failed hellos from
one source address within 60s, further connections are refused (lockout).
Heartbeat
→ { "type": "ping" }
← { "type": "pong", "ts": "..." }
The server also sends WebSocket protocol-level pings (aiohttp heartbeat=30).
Errors
Every error uses one shape with a machine-readable code:
{ "type": "error", "request_id": "r1",
"error": { "code": "conversation_not_found", "message": "Conversation not found" },
"ts": "..." }
Codes: unauthorized, auth_timeout, version_mismatch, bad_request,
invalid_json, unknown_type, not_found, conversation_not_found,
approval_not_found, clarify_not_found, too_large, rate_limited,
run_active, internal_error, not_implemented.
Limits: chat text ≤ 64,000 chars; inbound WS frame ≤ 2 MiB (violations get
too_large); conversation history fetch ≤ 500 messages.
Conversations
conversation.list
→ { "type": "conversation.list", "request_id": "r1" }
← { "type": "conversation.snapshot", "request_id": "r1",
"conversations": [
{ "conversation_id": "9f1c…", "name": "Project X",
"session_id": "20260902_101112_ab12cd34", // Hermes session (may be null)
"last_active": "2026-09-02T17:44:01+00:00", // may be null
"source": "hermes" } ] }
conversation.create
→ { "type": "conversation.create", "name": "New chat", "request_id": "r2" }
← { "type": "conversation.created", "conversation_id": "a1b2…", "name": "New chat", "request_id": "r2" }
// plus, broadcast to all clients:
← { "type": "conversation.updated", "conversation_id": "a1b2…", "name": "New chat" }
name optional. Conversation IDs are server-generated 32-hex opaque strings.
conversation.open — load history (reconnect recovery)
→ { "type": "conversation.open", "conversation_id": "a1b2…", "limit": 200, "request_id": "r3" }
← { "type": "conversation.history", "conversation_id": "a1b2…", "request_id": "r3",
"messages": [
{ "message_id": "m12", "role": "user", "text": "hey", "ts": "…|null" },
{ "message_id": "m13", "role": "assistant", "text": "hi!", "ts": "…|null" } ],
"has_more": true,
"attachments": [],
"run": null,
"tools": [],
"approvals": [],
"clarifications": [] }
History is the authoritative Hermes transcript (role is always user or
assistant). It spans all durable reset and compression generations sharing
the Pheby conversation key. limit is capped at 500; the response returns the
newest page in chronological order. While has_more is true, request earlier
pages with "before_message_id": "m12", using the first message ID from the
previous page as an exclusive cursor. This is display history only; model
context can still be compressed. The other fields form a recoverable snapshot: unexpired
attachments, the active run (if any), latest structured tool states, and
pending approval/clarification requests. On reconnect, replace local state
with this snapshot, then consume new live events. Unknown conversation →
conversation_not_found error.
conversation.rename
→ { "type": "conversation.rename", "conversation_id": "a1b2…", "name": "Renamed", "request_id": "r4" }
← { "type": "conversation.renamed", "conversation_id": "a1b2…", "name": "Renamed", "request_id": "r4" }
← { "type": "conversation.renamed", "conversation_id": "a1b2…", "name": "Renamed" } // broadcast
conversation.delete
→ { "type": "conversation.delete", "conversation_id": "a1b2…", "request_id": "r5" }
← { "type": "conversation.deleted", "conversation_id": "a1b2…", "request_id": "r5" }
Deletes the Hermes session transcript and the routing entry. Attachments belonging to the conversation age out on their own 7-day schedule.
Not supported by design: message editing, per-message deletion, regeneration, edit-and-resend. If you need to "undo", send a correction message (the agent sees the whole transcript).
Chat & streaming
chat.send
→ { "type": "chat.send", "conversation_id": "a1b2…", "text": "What's the weather?", "request_id": "r6" }
attachment_ids is an optional list of up to 10 distinct inbound upload IDs
from POST /attachments in the same conversation. The adapter rejects
unknown, already-sent, or cross-conversation IDs. Upload first, then include
all IDs in the single chat.send; a rejected active run leaves them retryable.
The text field is still required and nonblank (for file-only sends, provide
a short caption). A successful send anchors the attachments to the user turn
and broadcasts attachment.added. Small text files are included in agent
context; other files remain available to the agent as local media paths.
Server → Client run lifecycle
← { "type": "run.accepted", "conversation_id": "a1b2…", "run_id": "8c1f…", "request_id": "r6" }
← { "type": "message.start", "conversation_id": "a1b2…", "run_id": "8c1f…", "message_id": "draft-8c1f…" }
While the agent streams, the server pushes cumulative draft text (the client can simply replace the bubble's text each time — no delta stitching):
← { "type": "message.delta", "conversation_id": "a1b2…", "message_id": "draft-8c1f…",
"text": "It's currently 27°C…", "ts": "..." }
Completion (final text supersedes the draft — render the final, drop the draft):
← { "type": "message.complete", "conversation_id": "a1b2…",
"message_id": "draft-8c1f…", "text": "…full final answer…", "ts": "..." }
message.complete with "kind": "notice" is a gateway lifecycle/status
notice rather than conversation content — render or ignore.
Run end:
← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…",
"status": "completed" | "cancelled" | "failed",
"error": "only on failure" }
State machine per assistant turn:
run.accepted → message.start → (message.delta)* → message.complete → run.finished.
A non-streaming turn may skip message.delta; message.start is always sent
when the run is accepted. Only one run may be active per conversation;
another chat.send receives run_active. Never infer state from text.
run.cancel — stop an active run
→ { "type": "run.cancel", "conversation_id": "a1b2…", "run_id": "8c1f…", "request_id": "r7" }
← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…", "status": "cancelled" }
Cancellation uses Hermes's supported interrupt mechanism (agent interrupt +
run-generation invalidation) — the conversation stays consistent and
resumable. A stale run_id, missing run, or run that Hermes can no longer
interrupt returns not_found; success emits exactly one cancelled event.
Tool events (structured — never chat text)
Tool activity arrives as tool.event messages, completely separate from
message.* chat content. The client renders them as compact tool components
(ChatGPT-style) attached to the assistant turn.
{ "type": "tool.event",
"conversation_id": "a1b2…",
"run_id": "8c1f…",
"tool_call_id": "t-1a2b3c4d5e6f",
"tool_name": "web_search",
"status": "running" | "completed" | "failed" | "cancelled",
"args_redacted": { "query": "cats" }, // only on "running"; secret-looking keys redacted
"duration_ms": 1234, // only on completion/failure (may be null)
"error": "only on failed, truncated", // only on "failed"
"ts": "..." }
Correlate running → completed/failed by tool_call_id. Both events use
Hermes's authoritative call ID from the pre/post tool hooks and include the
active run_id when one exists. Secret-looking argument keys are redacted
recursively before leaving the server.
No fake "Searching the web…" text is ever injected into message.* events.
Approvals
When Hermes pauses for a human decision on a dangerous action:
{ "type": "approval.request",
"approval_id": "3d4e5f6070a1",
"session_key": "agent:main:pheby:dm:a1b2…",
"conversation_id": "a1b2…",
"run_id": "8c1f…",
"command": "rm -rf /tmp/build-output",
"description": "Destructive shell command (rm -rf)",
"choices": ["once", "session", "always", "deny"],
"ts": "..." }
Respond:
→ { "type": "approval.respond", "approval_id": "3d4e5f6070a1",
"choice": "once" | "session" | "always" | "deny",
"reason": "optional free text with deny", "request_id": "r8" }
← { "type": "approval.resolved", "approval_id": "3d4e5f6070a1",
"choice": "once", "accepted": true, "request_id": "r8" }
The single resolution event is broadcast to every connected client; the
responding client's request_id is included on that broadcast.
Choices map to Hermes semantics: once (approve this action), session
(approve pattern for this conversation), always (also persist), deny
(decline; the agent is told NOT to retry). Unknown/stale ID →
approval_not_found error. Hermes itself fails the approval closed after its
own timeout, so a silently-closed socket can't leave a zombie gate.
Clarifications / choices
{ "type": "clarify.request",
"clarify_id": "c1a2b3d4e5",
"session_key": "agent:main:pheby:dm:a1b2…",
"conversation_id": "a1b2…",
"run_id": "8c1f…",
"question": "Deploy to staging or production?",
"choices": ["staging", "production"], // null ⇒ free text only
"allow_free_text": true,
"ts": "..." }
Respond (either a choice value or free text):
→ { "type": "clarify.respond", "clarify_id": "c1a2b3d4e5", "response": "production", "request_id": "r9" }
← { "type": "clarify.resolved", "clarify_id": "c1a2b3d4e5",
"accepted": true, "request_id": "r9" }
The single resolution event is broadcast. A stale/timed-out ID instead gets
clarify_not_found. Always render an "Other" affordance — Hermes
clarifications accept free text.
Auto-approve (YOLO)
YOLO is always conversation-scoped. It skips ordinary dangerous-command approval prompts for that Hermes session, while Hermes hard deny rules and hardline safety floors still apply.
→ { "type": "yolo.current", "conversation_id": "a1b2…", "request_id": "r15" }
← { "type": "yolo.snapshot", "enabled": false, "scope": "conversation",
"conversation_id": "a1b2…", "request_id": "r15" }
→ { "type": "yolo.set", "conversation_id": "a1b2…", "enabled": true,
"request_id": "r16" }
← { "type": "yolo.changed", "enabled": true, "scope": "conversation",
"conversation_id": "a1b2…", "request_id": "r16" }
// plus broadcast of yolo.changed (without request_id) to all clients
enabled must be a JSON boolean. Unknown conversations return
conversation_not_found. The adapter also persists the flag in Hermes session
metadata so it can be restored after the gateway restarts.
Attachments (both directions)
Client → agent uploads
POST /attachments?conversation_id=a1b2…&filename=notes.md
Authorization: Bearer <PHEBY_SECRET>
Content-Type: text/markdown
# Raw file bytes, not multipart or JSON
Returns 201 {"attachment": {"attachment_id": "<32 hex chars>", ...}}.
The descriptor contains direction: "inbound", message_id: null, and the
same fields as agent deliverables below. An upload is not broadcast or
listed in conversation history until chat.send successfully claims it;
an unused upload expires with normal retention. Limit: 64 MiB per file.
Missing credentials → 401; bad conversation ID or empty body → 400;
oversize body → 413. The file picker may select up to 10 files per message.
The server keeps only registered copies in adapter-owned storage and never
trusts a client-provided path. Markdown and other small text/* files
(up to 100 KiB) are inlined into the agent's turn; the conversation history
returns just the user's original message text.
Agent → client deliverables
When the agent produces a file (image, document, audio, video… via Hermes's
normal MEDIA: deliverable pipeline), the server copies it into
adapter-managed storage and broadcasts:
{ "type": "attachment.added",
"conversation_id": "a1b2…",
"attachment": {
"attachment_id": "e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0",
"filename": "report.pdf",
"mime_type": "application/pdf",
"size": 48213,
"kind": "image" | "voice" | "video" | "audio" | "document",
"inline_image": false,
"direction": "outbound",
"conversation_id": "a1b2…",
"message_id": "draft-8c1f…" | null,
"created_at": "2026-09-02T18:30:00+00:00",
"expires_at": "2026-09-09T18:30:00+00:00", // null when retention=0
"download_path": "/attachments/e5f6…" },
"ts": "..." }
Download over HTTPS (authenticated — same PHEBY_SECRET):
GET {download_path}
Authorization: Bearer <PHEBY_SECRET> (or ApiKey <secret>, or X-Pheby-Secret: <secret>)
inline_image: true→kind == "image", safe for an inline preview (BitmapFactory/AsyncImagewith the same authenticated GET).- Everything else: download and open as an Android document.
- Unknown / expired / malformed ID →
404with{"error":{"code":"not_found","message":"Attachment unavailable"}}— no implementation details. - The client can never request arbitrary files — only registered attachment IDs resolve.
- Retention: adapter copies expire after 7 days (configurable) and are
deleted by an hourly cleanup. Original files the agent produced elsewhere
on the host are never touched.
expires_attells the client when to stop offering the download.
Models
List providers + models
→ { "type": "models.list", "conversation_id": "a1b2…", "request_id": "r10" }
← { "type": "models.snapshot", "request_id": "r10",
"providers": [
{ "slug": "openrouter", "name": "OpenRouter", "is_current": true,
"models": ["z-ai/glm-5.3-flash", "anthropic/claude-sonnet-4", "…"],
"total_models": 42 } ],
"current_model": "z-ai/glm-5.3-flash",
"current_provider": "openrouter",
"supported_reasoning_efforts": ["minimal","low","medium","high","xhigh","max","ultra"],
"ts": "..." }
Lists come from Hermes's own credential-aware picker data — nothing is
hardcoded. Models are exactly what the configured providers expose. The
optional conversation_id makes current_model, current_provider, and
scope reflect that conversation's override.
Read / change current model
→ { "type": "models.current", "conversation_id": "a1b2…", "request_id": "r11" }
← { "type": "model.current", "request_id": "r11", "model": "z-ai/glm-5.3-flash",
"provider": "openrouter", "scope": "conversation",
"conversation_id": "a1b2…", "ts": "..." }
→ { "type": "model.set", "model": "anthropic/claude-sonnet-4",
"provider": "anthropic", // optional
"conversation_id": "a1b2…", // present ⇒ session-scoped override
"request_id": "r12" }
← { "type": "model.changed", "model": "anthropic/claude-sonnet-4",
"provider": "anthropic", "scope": "conversation" | "global", "request_id": "r12" }
// plus broadcast of model.changed (without request_id) to all clients
Omit conversation_id ⇒ the change is persisted globally (Hermes
model.default). A model.changed with scope:"global" tells every open
conversation the default moved.
Reasoning effort
→ { "type": "reasoning.current", "conversation_id": "a1b2…", "request_id": "r13" }
← { "type": "reasoning.snapshot", "request_id": "r13",
"effort": "medium", // current effective effort (may be null = provider default)
"enabled": true, // false ⇒ thinking disabled
"scope": "conversation", "conversation_id": "a1b2…",
"supported_efforts": ["none","minimal","low","medium","high","xhigh","max","ultra"],
"ts": "..." }
→ { "type": "reasoning.set", "effort": "high", "conversation_id": "a1b2…", "request_id": "r14" }
← { "type": "reasoning.changed", "effort": "high", "scope": "conversation", "request_id": "r14" }
effort: "none" disables thinking. Invalid values → bad_request. Scope
rules mirror model.set (with conversation_id ⇒ session override; without ⇒
global agent.reasoning_effort). Capability note: Hermes knows whether
a model supports reasoning (models.dev metadata) but does not expose a
per-provider enum of valid effort values; the listed levels are Hermes's
canonical set — unsupported levels on a given provider surface as a provider
error on the next turn, not at set time. This is a documented Hermes
limitation, not a Pheby guess.
Unsolicited messages & reconnect behavior
The WebSocket stays connected; any Hermes-originated output destined for the
Pheby platform (scheduled/cron deliveries, background completions,
notifications) is pushed as normal message.* / run.* events even when it
is not a reply to your last request. Reconnection procedure for clients:
- Reconnect WS, redo
hello. - Re-
conversation.openthe conversations you show; replace local history, attachments, run/tool state, and pending decisions with its authoritative snapshot. - Re-
models.current/reasoning.currentif those views are visible. - Live events continue from there.
No external push service exists (no FCM); Android notification behavior is the client's responsibility while the socket is down.