Files
pheby-hermes-platform-adapter/docs/PROTOCOL.md
T

18 KiB

Pheby Protocol v1 — Specification

JSON messages over WebSocket, plus authenticated HTTPS endpoints for attachment uploads and downloads. Every message (both directions) carries a "type". Client→server requests MAY carry a "request_id" (any client-chosen string); the direct reply echoes it. Server→client events are broadcast to all authenticated connections (single-user app — typically one client).

Timestamps ("ts") are ISO-8601 UTC. All IDs are opaque strings — the client never constructs meaning from them and never sees server filesystem paths.

Handshake

Connect to wss://<host>/ws (behind Caddy; the plugin itself is plain ws://127.0.0.1:8620/ws). The first client frame must be hello within 10 seconds, or the server closes the socket (auth_timeout).

Client → Server hello

{
  "type": "hello",
  "secret": "<PHEBY_SECRET>",
  "protocol_version": 1,
  "request_id": "optional"
}

Server → Client ready (success)

{ "type": "ready", "protocol_version": 1, "server": "pheby",
  "ts": "2026-09-02T18:00:00+00:00" }

Failure: the server replies with an error event (unauthorized, auth_timeout, or version_mismatch) and closes. After 5 failed hellos from one source address within 60s, further connections are refused (lockout).

Heartbeat

→ { "type": "ping" }
← { "type": "pong", "ts": "..." }

The server also sends WebSocket protocol-level pings (aiohttp heartbeat=30).

Errors

Every error uses one shape with a machine-readable code:

{ "type": "error", "request_id": "r1",
  "error": { "code": "conversation_not_found", "message": "Conversation not found" },
  "ts": "..." }

Codes: unauthorized, auth_timeout, version_mismatch, bad_request, invalid_json, unknown_type, not_found, conversation_not_found, approval_not_found, clarify_not_found, too_large, rate_limited, run_active, internal_error, not_implemented.

Limits: chat text ≤ 64,000 chars; inbound WS frame ≤ 2 MiB (violations get too_large); conversation history fetch ≤ 500 messages.


Conversations

conversation.list

→ { "type": "conversation.list", "request_id": "r1" }
← { "type": "conversation.snapshot", "request_id": "r1",
    "conversations": [
      { "conversation_id": "9f1c…", "name": "Project X",
        "session_id": "20260902_101112_ab12cd34",     // Hermes session (may be null)
        "last_active": "2026-09-02T17:44:01+00:00",   // may be null
        "source": "hermes" } ] }

conversation.create

→ { "type": "conversation.create", "name": "New chat", "request_id": "r2" }
← { "type": "conversation.created", "conversation_id": "a1b2…", "name": "New chat", "request_id": "r2" }
// plus, broadcast to all clients:
← { "type": "conversation.updated", "conversation_id": "a1b2…", "name": "New chat" }

name optional. Conversation IDs are server-generated 32-hex opaque strings.

conversation.open — load history (reconnect recovery)

→ { "type": "conversation.open", "conversation_id": "a1b2…", "limit": 200, "request_id": "r3" }
← { "type": "conversation.history", "conversation_id": "a1b2…", "request_id": "r3",
    "messages": [
      { "message_id": "m12", "role": "user",      "text": "hey",  "ts": "…|null" },
      { "message_id": "m13", "role": "assistant", "text": "hi!",  "ts": "…|null" } ],
    "attachments": [],
    "run": null,
    "tools": [],
    "approvals": [],
    "clarifications": [] }

History is the authoritative Hermes transcript (role is always user or assistant). The other fields form a recoverable snapshot: unexpired attachments, the active run (if any), latest structured tool states, and pending approval/clarification requests. On reconnect, replace local state with this snapshot, then consume new live events. Unknown conversation → conversation_not_found error.

conversation.rename

→ { "type": "conversation.rename", "conversation_id": "a1b2…", "name": "Renamed", "request_id": "r4" }
← { "type": "conversation.renamed", "conversation_id": "a1b2…", "name": "Renamed", "request_id": "r4" }
← { "type": "conversation.renamed", "conversation_id": "a1b2…", "name": "Renamed" }   // broadcast

conversation.delete

→ { "type": "conversation.delete", "conversation_id": "a1b2…", "request_id": "r5" }
← { "type": "conversation.deleted", "conversation_id": "a1b2…", "request_id": "r5" }

Deletes the Hermes session transcript and the routing entry. Attachments belonging to the conversation age out on their own 7-day schedule.

Not supported by design: message editing, per-message deletion, regeneration, edit-and-resend. If you need to "undo", send a correction message (the agent sees the whole transcript).


Chat & streaming

chat.send

→ { "type": "chat.send", "conversation_id": "a1b2…", "text": "What's the weather?", "request_id": "r6" }

attachment_ids is an optional list of up to 10 distinct inbound upload IDs from POST /attachments in the same conversation. The adapter rejects unknown, already-sent, or cross-conversation IDs. Upload first, then include all IDs in the single chat.send; a rejected active run leaves them retryable. The text field is still required and nonblank (for file-only sends, provide a short caption). A successful send anchors the attachments to the user turn and broadcasts attachment.added. Small text files are included in agent context; other files remain available to the agent as local media paths.

Server → Client run lifecycle

← { "type": "run.accepted", "conversation_id": "a1b2…", "run_id": "8c1f…", "request_id": "r6" }
← { "type": "message.start", "conversation_id": "a1b2…", "run_id": "8c1f…", "message_id": "draft-8c1f…" }

While the agent streams, the server pushes cumulative draft text (the client can simply replace the bubble's text each time — no delta stitching):

← { "type": "message.delta", "conversation_id": "a1b2…", "message_id": "draft-8c1f…",
    "text": "It's currently 27°C…", "ts": "..." }

Completion (final text supersedes the draft — render the final, drop the draft):

← { "type": "message.complete", "conversation_id": "a1b2…",
    "message_id": "draft-8c1f…", "text": "…full final answer…", "ts": "..." }

message.complete with "kind": "notice" is a gateway lifecycle/status notice rather than conversation content — render or ignore.

Run end:

← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…",
    "status": "completed" | "cancelled" | "failed",
    "error": "only on failure" }

State machine per assistant turn: run.accepted → message.start → (message.delta)* → message.complete → run.finished. A non-streaming turn may skip message.delta; message.start is always sent when the run is accepted. Only one run may be active per conversation; another chat.send receives run_active. Never infer state from text.

run.cancel — stop an active run

→ { "type": "run.cancel", "conversation_id": "a1b2…", "run_id": "8c1f…", "request_id": "r7" }
← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…", "status": "cancelled" }

Cancellation uses Hermes's supported interrupt mechanism (agent interrupt + run-generation invalidation) — the conversation stays consistent and resumable. A stale run_id, missing run, or run that Hermes can no longer interrupt returns not_found; success emits exactly one cancelled event.


Tool events (structured — never chat text)

Tool activity arrives as tool.event messages, completely separate from message.* chat content. The client renders them as compact tool components (ChatGPT-style) attached to the assistant turn.

{ "type": "tool.event",
  "conversation_id": "a1b2…",
  "run_id": "8c1f…",
  "tool_call_id": "t-1a2b3c4d5e6f",
  "tool_name": "web_search",
  "status": "running" | "completed" | "failed" | "cancelled",
  "args_redacted": { "query": "cats" },          // only on "running"; secret-looking keys redacted
  "duration_ms": 1234,                            // only on completion/failure (may be null)
  "error": "only on failed, truncated",           // only on "failed"
  "ts": "..." }

Correlate running → completed/failed by tool_call_id. Both events use Hermes's authoritative call ID from the pre/post tool hooks and include the active run_id when one exists. Secret-looking argument keys are redacted recursively before leaving the server.

No fake "Searching the web…" text is ever injected into message.* events.


Approvals

When Hermes pauses for a human decision on a dangerous action:

{ "type": "approval.request",
  "approval_id": "3d4e5f6070a1",
  "session_key": "agent:main:pheby:dm:a1b2…",
  "conversation_id": "a1b2…",
  "run_id": "8c1f…",
  "command": "rm -rf /tmp/build-output",
  "description": "Destructive shell command (rm -rf)",
  "choices": ["once", "session", "always", "deny"],
  "ts": "..." }

Respond:

→ { "type": "approval.respond", "approval_id": "3d4e5f6070a1",
    "choice": "once" | "session" | "always" | "deny",
    "reason": "optional free text with deny", "request_id": "r8" }
← { "type": "approval.resolved", "approval_id": "3d4e5f6070a1",
    "choice": "once", "accepted": true, "request_id": "r8" }

The single resolution event is broadcast to every connected client; the responding client's request_id is included on that broadcast.

Choices map to Hermes semantics: once (approve this action), session (approve pattern for this conversation), always (also persist), deny (decline; the agent is told NOT to retry). Unknown/stale ID → approval_not_found error. Hermes itself fails the approval closed after its own timeout, so a silently-closed socket can't leave a zombie gate.


Clarifications / choices

{ "type": "clarify.request",
  "clarify_id": "c1a2b3d4e5",
  "session_key": "agent:main:pheby:dm:a1b2…",
  "conversation_id": "a1b2…",
  "run_id": "8c1f…",
  "question": "Deploy to staging or production?",
  "choices": ["staging", "production"],      // null ⇒ free text only
  "allow_free_text": true,
  "ts": "..." }

Respond (either a choice value or free text):

→ { "type": "clarify.respond", "clarify_id": "c1a2b3d4e5", "response": "production", "request_id": "r9" }
← { "type": "clarify.resolved", "clarify_id": "c1a2b3d4e5",
    "accepted": true, "request_id": "r9" }

The single resolution event is broadcast. A stale/timed-out ID instead gets clarify_not_found. Always render an "Other" affordance — Hermes clarifications accept free text.


Auto-approve (YOLO)

YOLO is always conversation-scoped. It skips ordinary dangerous-command approval prompts for that Hermes session, while Hermes hard deny rules and hardline safety floors still apply.

→ { "type": "yolo.current", "conversation_id": "a1b2…", "request_id": "r15" }
← { "type": "yolo.snapshot", "enabled": false, "scope": "conversation",
    "conversation_id": "a1b2…", "request_id": "r15" }

→ { "type": "yolo.set", "conversation_id": "a1b2…", "enabled": true,
    "request_id": "r16" }
← { "type": "yolo.changed", "enabled": true, "scope": "conversation",
    "conversation_id": "a1b2…", "request_id": "r16" }
// plus broadcast of yolo.changed (without request_id) to all clients

enabled must be a JSON boolean. Unknown conversations return conversation_not_found. The adapter also persists the flag in Hermes session metadata so it can be restored after the gateway restarts.


Attachments (both directions)

Client → agent uploads

POST /attachments?conversation_id=a1b2…&filename=notes.md
Authorization: Bearer <PHEBY_SECRET>
Content-Type: text/markdown

# Raw file bytes, not multipart or JSON

Returns 201 {"attachment": {"attachment_id": "<32 hex chars>", ...}}. The descriptor contains direction: "inbound", message_id: null, and the same fields as agent deliverables below. An upload is not broadcast or listed in conversation history until chat.send successfully claims it; an unused upload expires with normal retention. Limit: 64 MiB per file. Missing credentials → 401; bad conversation ID or empty body → 400; oversize body → 413. The file picker may select up to 10 files per message. The server keeps only registered copies in adapter-owned storage and never trusts a client-provided path. Markdown and other small text/* files (up to 100 KiB) are inlined into the agent's turn; the conversation history returns just the user's original message text.

Agent → client deliverables

When the agent produces a file (image, document, audio, video… via Hermes's normal MEDIA: deliverable pipeline), the server copies it into adapter-managed storage and broadcasts:

{ "type": "attachment.added",
  "conversation_id": "a1b2…",
  "attachment": {
    "attachment_id": "e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0",
    "filename": "report.pdf",
    "mime_type": "application/pdf",
    "size": 48213,
    "kind": "image" | "voice" | "video" | "audio" | "document",
    "inline_image": false,
    "direction": "outbound",
    "conversation_id": "a1b2…",
    "message_id": "draft-8c1f…" | null,
    "created_at": "2026-09-02T18:30:00+00:00",
    "expires_at": "2026-09-09T18:30:00+00:00",     // null when retention=0
    "download_path": "/attachments/e5f6…" },
  "ts": "..." }

Download over HTTPS (authenticated — same PHEBY_SECRET):

GET {download_path}
Authorization: Bearer <PHEBY_SECRET>       (or ApiKey <secret>, or X-Pheby-Secret: <secret>)
  • inline_image: true → kind == "image", safe for an inline preview (BitmapFactory / AsyncImage with the same authenticated GET).
  • Everything else: download and open as an Android document.
  • Unknown / expired / malformed ID → 404 with {"error":{"code":"not_found","message":"Attachment unavailable"}} — no implementation details.
  • The client can never request arbitrary files — only registered attachment IDs resolve.
  • Retention: adapter copies expire after 7 days (configurable) and are deleted by an hourly cleanup. Original files the agent produced elsewhere on the host are never touched. expires_at tells the client when to stop offering the download.

Models

List providers + models

→ { "type": "models.list", "conversation_id": "a1b2…", "request_id": "r10" }
← { "type": "models.snapshot", "request_id": "r10",
    "providers": [
      { "slug": "openrouter", "name": "OpenRouter", "is_current": true,
        "models": ["z-ai/glm-5.3-flash", "anthropic/claude-sonnet-4", "…"],
        "total_models": 42 } ],
    "current_model": "z-ai/glm-5.3-flash",
    "current_provider": "openrouter",
    "supported_reasoning_efforts": ["minimal","low","medium","high","xhigh","max","ultra"],
    "ts": "..." }

Lists come from Hermes's own credential-aware picker data — nothing is hardcoded. Models are exactly what the configured providers expose. The optional conversation_id makes current_model, current_provider, and scope reflect that conversation's override.

Read / change current model

→ { "type": "models.current", "conversation_id": "a1b2…", "request_id": "r11" }
← { "type": "model.current", "request_id": "r11", "model": "z-ai/glm-5.3-flash",
    "provider": "openrouter", "scope": "conversation",
    "conversation_id": "a1b2…", "ts": "..." }

→ { "type": "model.set", "model": "anthropic/claude-sonnet-4",
    "provider": "anthropic",                  // optional
    "conversation_id": "a1b2…",               // present ⇒ session-scoped override
    "request_id": "r12" }
← { "type": "model.changed", "model": "anthropic/claude-sonnet-4",
    "provider": "anthropic", "scope": "conversation" | "global", "request_id": "r12" }
// plus broadcast of model.changed (without request_id) to all clients

Omit conversation_id ⇒ the change is persisted globally (Hermes model.default). A model.changed with scope:"global" tells every open conversation the default moved.


Reasoning effort

→ { "type": "reasoning.current", "conversation_id": "a1b2…", "request_id": "r13" }
← { "type": "reasoning.snapshot", "request_id": "r13",
    "effort": "medium",                       // current effective effort (may be null = provider default)
    "enabled": true,                          // false ⇒ thinking disabled
    "scope": "conversation", "conversation_id": "a1b2…",
    "supported_efforts": ["none","minimal","low","medium","high","xhigh","max","ultra"],
    "ts": "..." }

→ { "type": "reasoning.set", "effort": "high", "conversation_id": "a1b2…", "request_id": "r14" }
← { "type": "reasoning.changed", "effort": "high", "scope": "conversation", "request_id": "r14" }

effort: "none" disables thinking. Invalid values → bad_request. Scope rules mirror model.set (with conversation_id ⇒ session override; without ⇒ global agent.reasoning_effort). Capability note: Hermes knows whether a model supports reasoning (models.dev metadata) but does not expose a per-provider enum of valid effort values; the listed levels are Hermes's canonical set — unsupported levels on a given provider surface as a provider error on the next turn, not at set time. This is a documented Hermes limitation, not a Pheby guess.


Unsolicited messages & reconnect behavior

The WebSocket stays connected; any Hermes-originated output destined for the Pheby platform (scheduled/cron deliveries, background completions, notifications) is pushed as normal message.* / run.* events even when it is not a reply to your last request. Reconnection procedure for clients:

  1. Reconnect WS, redo hello.
  2. Re-conversation.open the conversations you show; replace local history, attachments, run/tool state, and pending decisions with its authoritative snapshot.
  3. Re-models.current / reasoning.current if those views are visible.
  4. Live events continue from there.

No external push service exists (no FCM); Android notification behavior is the client's responsibility while the socket is down.