492 lines
18 KiB
Markdown
492 lines
18 KiB
Markdown
# Pheby Protocol v1 — Specification
|
|
|
|
JSON messages over WebSocket, plus authenticated HTTPS endpoints for
|
|
attachment uploads and downloads. Every message (both directions) carries a `"type"`.
|
|
Client→server requests MAY carry a `"request_id"` (any client-chosen string);
|
|
the direct reply echoes it. Server→client events are broadcast to all
|
|
authenticated connections (single-user app — typically one client).
|
|
|
|
Timestamps (`"ts"`) are ISO-8601 UTC. All IDs are opaque strings — the client
|
|
never constructs meaning from them and never sees server filesystem paths.
|
|
|
|
## Handshake
|
|
|
|
Connect to `wss://<host>/ws` (behind Caddy; the plugin itself is plain
|
|
`ws://127.0.0.1:8620/ws`). The **first** client frame must be `hello` within
|
|
10 seconds, or the server closes the socket (`auth_timeout`).
|
|
|
|
### Client → Server `hello`
|
|
|
|
```json
|
|
{
|
|
"type": "hello",
|
|
"secret": "<PHEBY_SECRET>",
|
|
"protocol_version": 1,
|
|
"request_id": "optional"
|
|
}
|
|
```
|
|
|
|
### Server → Client `ready` (success)
|
|
|
|
```json
|
|
{ "type": "ready", "protocol_version": 1, "server": "pheby",
|
|
"ts": "2026-09-02T18:00:00+00:00" }
|
|
```
|
|
|
|
Failure: the server replies with an `error` event (`unauthorized`,
|
|
`auth_timeout`, or `version_mismatch`) and closes. After 5 failed hellos from
|
|
one source address within 60s, further connections are refused (lockout).
|
|
|
|
### Heartbeat
|
|
|
|
```json
|
|
→ { "type": "ping" }
|
|
← { "type": "pong", "ts": "..." }
|
|
```
|
|
|
|
The server also sends WebSocket protocol-level pings (aiohttp `heartbeat=30`).
|
|
|
|
## Errors
|
|
|
|
Every error uses one shape with a machine-readable code:
|
|
|
|
```json
|
|
{ "type": "error", "request_id": "r1",
|
|
"error": { "code": "conversation_not_found", "message": "Conversation not found" },
|
|
"ts": "..." }
|
|
```
|
|
|
|
Codes: `unauthorized`, `auth_timeout`, `version_mismatch`, `bad_request`,
|
|
`invalid_json`, `unknown_type`, `not_found`, `conversation_not_found`,
|
|
`approval_not_found`, `clarify_not_found`, `too_large`, `rate_limited`,
|
|
`run_active`, `internal_error`, `not_implemented`.
|
|
|
|
Limits: chat text ≤ 64,000 chars; inbound WS frame ≤ 2 MiB (violations get
|
|
`too_large`); conversation history fetch ≤ 500 messages.
|
|
|
|
---
|
|
|
|
## Conversations
|
|
|
|
### `conversation.list`
|
|
|
|
```json
|
|
→ { "type": "conversation.list", "request_id": "r1" }
|
|
← { "type": "conversation.snapshot", "request_id": "r1",
|
|
"conversations": [
|
|
{ "conversation_id": "9f1c…", "name": "Project X",
|
|
"session_id": "20260902_101112_ab12cd34", // Hermes session (may be null)
|
|
"last_active": "2026-09-02T17:44:01+00:00", // may be null
|
|
"source": "hermes" } ] }
|
|
```
|
|
|
|
### `conversation.create`
|
|
|
|
```json
|
|
→ { "type": "conversation.create", "name": "New chat", "request_id": "r2" }
|
|
← { "type": "conversation.created", "conversation_id": "a1b2…", "name": "New chat", "request_id": "r2" }
|
|
// plus, broadcast to all clients:
|
|
← { "type": "conversation.updated", "conversation_id": "a1b2…", "name": "New chat" }
|
|
```
|
|
|
|
`name` optional. Conversation IDs are server-generated 32-hex opaque strings.
|
|
|
|
### `conversation.open` — load history (reconnect recovery)
|
|
|
|
```json
|
|
→ { "type": "conversation.open", "conversation_id": "a1b2…", "limit": 200, "request_id": "r3" }
|
|
← { "type": "conversation.history", "conversation_id": "a1b2…", "request_id": "r3",
|
|
"messages": [
|
|
{ "message_id": "m12", "role": "user", "text": "hey", "ts": "…|null" },
|
|
{ "message_id": "m13", "role": "assistant", "text": "hi!", "ts": "…|null" } ],
|
|
"has_more": true,
|
|
"attachments": [],
|
|
"run": null,
|
|
"tools": [],
|
|
"approvals": [],
|
|
"clarifications": [] }
|
|
```
|
|
|
|
History is the authoritative Hermes transcript (`role` is always `user` or
|
|
`assistant`). It spans all durable reset and compression generations sharing
|
|
the Pheby conversation key. `limit` is capped at 500; the response returns the
|
|
newest page in chronological order. While `has_more` is true, request earlier
|
|
pages with `"before_message_id": "m12"`, using the first message ID from the
|
|
previous page as an exclusive cursor. This is display history only; model
|
|
context can still be compressed. The other fields form a recoverable snapshot: unexpired
|
|
attachments, the active run (if any), latest structured tool states, and
|
|
pending approval/clarification requests. On reconnect, replace local state
|
|
with this snapshot, then consume new live events. Unknown conversation →
|
|
`conversation_not_found` error.
|
|
|
|
### `conversation.rename`
|
|
|
|
```json
|
|
→ { "type": "conversation.rename", "conversation_id": "a1b2…", "name": "Renamed", "request_id": "r4" }
|
|
← { "type": "conversation.renamed", "conversation_id": "a1b2…", "name": "Renamed", "request_id": "r4" }
|
|
← { "type": "conversation.renamed", "conversation_id": "a1b2…", "name": "Renamed" } // broadcast
|
|
```
|
|
|
|
### `conversation.delete`
|
|
|
|
```json
|
|
→ { "type": "conversation.delete", "conversation_id": "a1b2…", "request_id": "r5" }
|
|
← { "type": "conversation.deleted", "conversation_id": "a1b2…", "request_id": "r5" }
|
|
```
|
|
|
|
Deletes the Hermes session transcript and the routing entry. Attachments
|
|
belonging to the conversation age out on their own 7-day schedule.
|
|
|
|
**Not supported by design:** message editing, per-message deletion,
|
|
regeneration, edit-and-resend. If you need to "undo", send a correction
|
|
message (the agent sees the whole transcript).
|
|
|
|
---
|
|
|
|
## Chat & streaming
|
|
|
|
### `chat.send`
|
|
|
|
```json
|
|
→ { "type": "chat.send", "conversation_id": "a1b2…", "text": "What's the weather?", "request_id": "r6" }
|
|
```
|
|
|
|
`attachment_ids` is an optional list of up to 10 distinct inbound upload IDs
|
|
from `POST /attachments` in the **same conversation**. The adapter rejects
|
|
unknown, already-sent, or cross-conversation IDs. Upload first, then include
|
|
all IDs in the single `chat.send`; a rejected active run leaves them retryable.
|
|
The `text` field is still required and nonblank (for file-only sends, provide
|
|
a short caption). A successful send anchors the attachments to the user turn
|
|
and broadcasts `attachment.added`. Small text files are included in agent
|
|
context; other files remain available to the agent as local media paths.
|
|
|
|
### Server → Client run lifecycle
|
|
|
|
```json
|
|
← { "type": "run.accepted", "conversation_id": "a1b2…", "run_id": "8c1f…", "request_id": "r6" }
|
|
← { "type": "message.start", "conversation_id": "a1b2…", "run_id": "8c1f…", "message_id": "draft-8c1f…" }
|
|
```
|
|
|
|
While the agent streams, the server pushes **cumulative** draft text (the
|
|
client can simply replace the bubble's text each time — no delta stitching):
|
|
|
|
```json
|
|
← { "type": "message.delta", "conversation_id": "a1b2…", "message_id": "draft-8c1f…",
|
|
"text": "It's currently 27°C…", "ts": "..." }
|
|
```
|
|
|
|
Completion (final text supersedes the draft — render the final, drop the
|
|
draft):
|
|
|
|
```json
|
|
← { "type": "message.complete", "conversation_id": "a1b2…",
|
|
"message_id": "draft-8c1f…", "text": "…full final answer…", "ts": "..." }
|
|
```
|
|
|
|
`message.complete` with `"kind": "notice"` is a gateway lifecycle/status
|
|
notice rather than conversation content — render or ignore.
|
|
|
|
Run end:
|
|
|
|
```json
|
|
← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…",
|
|
"status": "completed" | "cancelled" | "failed",
|
|
"error": "only on failure" }
|
|
```
|
|
|
|
State machine per assistant turn:
|
|
`run.accepted → message.start → (message.delta)* → message.complete → run.finished`.
|
|
A non-streaming turn may skip `message.delta`; `message.start` is always sent
|
|
when the run is accepted. Only one run may be active per conversation;
|
|
another `chat.send` receives `run_active`. Never infer state from text.
|
|
|
|
### `run.cancel` — stop an active run
|
|
|
|
```json
|
|
→ { "type": "run.cancel", "conversation_id": "a1b2…", "run_id": "8c1f…", "request_id": "r7" }
|
|
← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…", "status": "cancelled" }
|
|
```
|
|
|
|
Cancellation uses Hermes's supported interrupt mechanism (agent interrupt +
|
|
run-generation invalidation) — the conversation stays consistent and
|
|
resumable. A stale `run_id`, missing run, or run that Hermes can no longer
|
|
interrupt returns `not_found`; success emits exactly one cancelled event.
|
|
|
|
---
|
|
|
|
## Tool events (structured — never chat text)
|
|
|
|
Tool activity arrives as `tool.event` messages, completely separate from
|
|
`message.*` chat content. The client renders them as compact tool components
|
|
(ChatGPT-style) attached to the assistant turn.
|
|
|
|
```json
|
|
{ "type": "tool.event",
|
|
"conversation_id": "a1b2…",
|
|
"run_id": "8c1f…",
|
|
"tool_call_id": "t-1a2b3c4d5e6f",
|
|
"tool_name": "web_search",
|
|
"status": "running" | "completed" | "failed" | "cancelled",
|
|
"args_redacted": { "query": "cats" }, // only on "running"; secret-looking keys redacted
|
|
"duration_ms": 1234, // only on completion/failure (may be null)
|
|
"error": "only on failed, truncated", // only on "failed"
|
|
"ts": "..." }
|
|
```
|
|
|
|
Correlate `running` → `completed`/`failed` by `tool_call_id`. Both events use
|
|
Hermes's authoritative call ID from the pre/post tool hooks and include the
|
|
active `run_id` when one exists. Secret-looking argument keys are redacted
|
|
recursively before leaving the server.
|
|
|
|
No fake "Searching the web…" text is ever injected into `message.*` events.
|
|
|
|
---
|
|
|
|
## Approvals
|
|
|
|
When Hermes pauses for a human decision on a dangerous action:
|
|
|
|
```json
|
|
{ "type": "approval.request",
|
|
"approval_id": "3d4e5f6070a1",
|
|
"session_key": "agent:main:pheby:dm:a1b2…",
|
|
"conversation_id": "a1b2…",
|
|
"run_id": "8c1f…",
|
|
"command": "rm -rf /tmp/build-output",
|
|
"description": "Destructive shell command (rm -rf)",
|
|
"choices": ["once", "session", "always", "deny"],
|
|
"ts": "..." }
|
|
```
|
|
|
|
Respond:
|
|
|
|
```json
|
|
→ { "type": "approval.respond", "approval_id": "3d4e5f6070a1",
|
|
"choice": "once" | "session" | "always" | "deny",
|
|
"reason": "optional free text with deny", "request_id": "r8" }
|
|
← { "type": "approval.resolved", "approval_id": "3d4e5f6070a1",
|
|
"choice": "once", "accepted": true, "request_id": "r8" }
|
|
```
|
|
|
|
The single resolution event is broadcast to every connected client; the
|
|
responding client's `request_id` is included on that broadcast.
|
|
|
|
Choices map to Hermes semantics: `once` (approve this action), `session`
|
|
(approve pattern for this conversation), `always` (also persist), `deny`
|
|
(decline; the agent is told NOT to retry). Unknown/stale ID →
|
|
`approval_not_found` error. Hermes itself fails the approval closed after its
|
|
own timeout, so a silently-closed socket can't leave a zombie gate.
|
|
|
|
---
|
|
|
|
## Clarifications / choices
|
|
|
|
```json
|
|
{ "type": "clarify.request",
|
|
"clarify_id": "c1a2b3d4e5",
|
|
"session_key": "agent:main:pheby:dm:a1b2…",
|
|
"conversation_id": "a1b2…",
|
|
"run_id": "8c1f…",
|
|
"question": "Deploy to staging or production?",
|
|
"choices": ["staging", "production"], // null ⇒ free text only
|
|
"allow_free_text": true,
|
|
"ts": "..." }
|
|
```
|
|
|
|
Respond (either a choice value or free text):
|
|
|
|
```json
|
|
→ { "type": "clarify.respond", "clarify_id": "c1a2b3d4e5", "response": "production", "request_id": "r9" }
|
|
← { "type": "clarify.resolved", "clarify_id": "c1a2b3d4e5",
|
|
"accepted": true, "request_id": "r9" }
|
|
```
|
|
|
|
The single resolution event is broadcast. A stale/timed-out ID instead gets
|
|
`clarify_not_found`. Always render an "Other" affordance — Hermes
|
|
clarifications accept free text.
|
|
|
|
---
|
|
|
|
## Auto-approve (YOLO)
|
|
|
|
YOLO is always conversation-scoped. It skips ordinary dangerous-command
|
|
approval prompts for that Hermes session, while Hermes hard deny rules and
|
|
hardline safety floors still apply.
|
|
|
|
```json
|
|
→ { "type": "yolo.current", "conversation_id": "a1b2…", "request_id": "r15" }
|
|
← { "type": "yolo.snapshot", "enabled": false, "scope": "conversation",
|
|
"conversation_id": "a1b2…", "request_id": "r15" }
|
|
|
|
→ { "type": "yolo.set", "conversation_id": "a1b2…", "enabled": true,
|
|
"request_id": "r16" }
|
|
← { "type": "yolo.changed", "enabled": true, "scope": "conversation",
|
|
"conversation_id": "a1b2…", "request_id": "r16" }
|
|
// plus broadcast of yolo.changed (without request_id) to all clients
|
|
```
|
|
|
|
`enabled` must be a JSON boolean. Unknown conversations return
|
|
`conversation_not_found`. The adapter also persists the flag in Hermes session
|
|
metadata so it can be restored after the gateway restarts.
|
|
|
|
---
|
|
|
|
## Attachments (both directions)
|
|
|
|
### Client → agent uploads
|
|
|
|
```
|
|
POST /attachments?conversation_id=a1b2…&filename=notes.md
|
|
Authorization: Bearer <PHEBY_SECRET>
|
|
Content-Type: text/markdown
|
|
|
|
# Raw file bytes, not multipart or JSON
|
|
```
|
|
|
|
Returns `201 {"attachment": {"attachment_id": "<32 hex chars>", ...}}`.
|
|
The descriptor contains `direction: "inbound"`, `message_id: null`, and the
|
|
same fields as agent deliverables below. An upload is **not** broadcast or
|
|
listed in conversation history until `chat.send` successfully claims it;
|
|
an unused upload expires with normal retention. Limit: 64 MiB per file.
|
|
Missing credentials → 401; bad conversation ID or empty body → 400;
|
|
oversize body → 413. The file picker may select up to 10 files per message.
|
|
The server keeps only registered copies in adapter-owned storage and never
|
|
trusts a client-provided path. Markdown and other small `text/*` files
|
|
(up to 100 KiB) are inlined into the agent's turn; the conversation history
|
|
returns just the user's original message text.
|
|
|
|
### Agent → client deliverables
|
|
|
|
When the agent produces a file (image, document, audio, video… via Hermes's
|
|
normal `MEDIA:` deliverable pipeline), the server copies it into
|
|
adapter-managed storage and broadcasts:
|
|
|
|
```json
|
|
{ "type": "attachment.added",
|
|
"conversation_id": "a1b2…",
|
|
"attachment": {
|
|
"attachment_id": "e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0",
|
|
"filename": "report.pdf",
|
|
"mime_type": "application/pdf",
|
|
"size": 48213,
|
|
"kind": "image" | "voice" | "video" | "audio" | "document",
|
|
"inline_image": false,
|
|
"direction": "outbound",
|
|
"conversation_id": "a1b2…",
|
|
"message_id": "draft-8c1f…" | null,
|
|
"created_at": "2026-09-02T18:30:00+00:00",
|
|
"expires_at": "2026-09-09T18:30:00+00:00", // null when retention=0
|
|
"download_path": "/attachments/e5f6…" },
|
|
"ts": "..." }
|
|
```
|
|
|
|
Download over **HTTPS** (authenticated — same `PHEBY_SECRET`):
|
|
|
|
```
|
|
GET {download_path}
|
|
Authorization: Bearer <PHEBY_SECRET> (or ApiKey <secret>, or X-Pheby-Secret: <secret>)
|
|
```
|
|
|
|
* `inline_image: true` → `kind == "image"`, safe for an inline preview
|
|
(`BitmapFactory` / `AsyncImage` with the same authenticated GET).
|
|
* Everything else: download and open as an Android document.
|
|
* Unknown / expired / malformed ID → `404` with
|
|
`{"error":{"code":"not_found","message":"Attachment unavailable"}}` —
|
|
no implementation details.
|
|
* **The client can never request arbitrary files** — only registered
|
|
attachment IDs resolve.
|
|
* **Retention:** adapter copies expire after 7 days (configurable) and are
|
|
deleted by an hourly cleanup. Original files the agent produced elsewhere
|
|
on the host are never touched. `expires_at` tells the client when to stop
|
|
offering the download.
|
|
|
|
---
|
|
|
|
## Models
|
|
|
|
### List providers + models
|
|
|
|
```json
|
|
→ { "type": "models.list", "conversation_id": "a1b2…", "request_id": "r10" }
|
|
← { "type": "models.snapshot", "request_id": "r10",
|
|
"providers": [
|
|
{ "slug": "openrouter", "name": "OpenRouter", "is_current": true,
|
|
"models": ["z-ai/glm-5.3-flash", "anthropic/claude-sonnet-4", "…"],
|
|
"total_models": 42 } ],
|
|
"current_model": "z-ai/glm-5.3-flash",
|
|
"current_provider": "openrouter",
|
|
"supported_reasoning_efforts": ["minimal","low","medium","high","xhigh","max","ultra"],
|
|
"ts": "..." }
|
|
```
|
|
|
|
Lists come from Hermes's own credential-aware picker data — nothing is
|
|
hardcoded. Models are exactly what the configured providers expose. The
|
|
optional `conversation_id` makes `current_model`, `current_provider`, and
|
|
`scope` reflect that conversation's override.
|
|
|
|
### Read / change current model
|
|
|
|
```json
|
|
→ { "type": "models.current", "conversation_id": "a1b2…", "request_id": "r11" }
|
|
← { "type": "model.current", "request_id": "r11", "model": "z-ai/glm-5.3-flash",
|
|
"provider": "openrouter", "scope": "conversation",
|
|
"conversation_id": "a1b2…", "ts": "..." }
|
|
|
|
→ { "type": "model.set", "model": "anthropic/claude-sonnet-4",
|
|
"provider": "anthropic", // optional
|
|
"conversation_id": "a1b2…", // present ⇒ session-scoped override
|
|
"request_id": "r12" }
|
|
← { "type": "model.changed", "model": "anthropic/claude-sonnet-4",
|
|
"provider": "anthropic", "scope": "conversation" | "global", "request_id": "r12" }
|
|
// plus broadcast of model.changed (without request_id) to all clients
|
|
```
|
|
|
|
Omit `conversation_id` ⇒ the change is persisted globally (Hermes
|
|
`model.default`). A `model.changed` with `scope:"global"` tells every open
|
|
conversation the default moved.
|
|
|
|
---
|
|
|
|
## Reasoning effort
|
|
|
|
```json
|
|
→ { "type": "reasoning.current", "conversation_id": "a1b2…", "request_id": "r13" }
|
|
← { "type": "reasoning.snapshot", "request_id": "r13",
|
|
"effort": "medium", // current effective effort (may be null = provider default)
|
|
"enabled": true, // false ⇒ thinking disabled
|
|
"scope": "conversation", "conversation_id": "a1b2…",
|
|
"supported_efforts": ["none","minimal","low","medium","high","xhigh","max","ultra"],
|
|
"ts": "..." }
|
|
|
|
→ { "type": "reasoning.set", "effort": "high", "conversation_id": "a1b2…", "request_id": "r14" }
|
|
← { "type": "reasoning.changed", "effort": "high", "scope": "conversation", "request_id": "r14" }
|
|
```
|
|
|
|
`effort: "none"` disables thinking. Invalid values → `bad_request`. Scope
|
|
rules mirror `model.set` (with `conversation_id` ⇒ session override; without ⇒
|
|
global `agent.reasoning_effort`). **Capability note:** Hermes knows *whether*
|
|
a model supports reasoning (models.dev metadata) but does not expose a
|
|
per-provider enum of valid effort values; the listed levels are Hermes's
|
|
canonical set — unsupported levels on a given provider surface as a provider
|
|
error on the next turn, not at set time. This is a documented Hermes
|
|
limitation, not a Pheby guess.
|
|
|
|
---
|
|
|
|
## Unsolicited messages & reconnect behavior
|
|
|
|
The WebSocket stays connected; any Hermes-originated output destined for the
|
|
Pheby platform (scheduled/cron deliveries, background completions,
|
|
notifications) is pushed as normal `message.*` / `run.*` events even when it
|
|
is not a reply to your last request. Reconnection procedure for clients:
|
|
|
|
1. Reconnect WS, redo `hello`.
|
|
2. Re-`conversation.open` the conversations you show; replace local history,
|
|
attachments, run/tool state, and pending decisions with its authoritative
|
|
snapshot.
|
|
3. Re-`models.current` / `reasoning.current` if those views are visible.
|
|
4. Live events continue from there.
|
|
|
|
No external push service exists (no FCM); Android notification behavior is
|
|
the client's responsibility while the socket is down.
|