Files

492 lines
18 KiB
Markdown

# Pheby Protocol v1 — Specification
JSON messages over WebSocket, plus authenticated HTTPS endpoints for
attachment uploads and downloads. Every message (both directions) carries a `"type"`.
Client→server requests MAY carry a `"request_id"` (any client-chosen string);
the direct reply echoes it. Server→client events are broadcast to all
authenticated connections (single-user app — typically one client).
Timestamps (`"ts"`) are ISO-8601 UTC. All IDs are opaque strings — the client
never constructs meaning from them and never sees server filesystem paths.
## Handshake
Connect to `wss://<host>/ws` (behind Caddy; the plugin itself is plain
`ws://127.0.0.1:8620/ws`). The **first** client frame must be `hello` within
10 seconds, or the server closes the socket (`auth_timeout`).
### Client → Server `hello`
```json
{
"type": "hello",
"secret": "<PHEBY_SECRET>",
"protocol_version": 1,
"request_id": "optional"
}
```
### Server → Client `ready` (success)
```json
{ "type": "ready", "protocol_version": 1, "server": "pheby",
"ts": "2026-09-02T18:00:00+00:00" }
```
Failure: the server replies with an `error` event (`unauthorized`,
`auth_timeout`, or `version_mismatch`) and closes. After 5 failed hellos from
one source address within 60s, further connections are refused (lockout).
### Heartbeat
```json
→ { "type": "ping" }
← { "type": "pong", "ts": "..." }
```
The server also sends WebSocket protocol-level pings (aiohttp `heartbeat=30`).
## Errors
Every error uses one shape with a machine-readable code:
```json
{ "type": "error", "request_id": "r1",
"error": { "code": "conversation_not_found", "message": "Conversation not found" },
"ts": "..." }
```
Codes: `unauthorized`, `auth_timeout`, `version_mismatch`, `bad_request`,
`invalid_json`, `unknown_type`, `not_found`, `conversation_not_found`,
`approval_not_found`, `clarify_not_found`, `too_large`, `rate_limited`,
`run_active`, `internal_error`, `not_implemented`.
Limits: chat text ≤ 64,000 chars; inbound WS frame ≤ 2 MiB (violations get
`too_large`); conversation history fetch ≤ 500 messages.
---
## Conversations
### `conversation.list`
```json
→ { "type": "conversation.list", "request_id": "r1" }
← { "type": "conversation.snapshot", "request_id": "r1",
"conversations": [
{ "conversation_id": "9f1c…", "name": "Project X",
"session_id": "20260902_101112_ab12cd34", // Hermes session (may be null)
"last_active": "2026-09-02T17:44:01+00:00", // may be null
"source": "hermes" } ] }
```
### `conversation.create`
```json
→ { "type": "conversation.create", "name": "New chat", "request_id": "r2" }
← { "type": "conversation.created", "conversation_id": "a1b2…", "name": "New chat", "request_id": "r2" }
// plus, broadcast to all clients:
← { "type": "conversation.updated", "conversation_id": "a1b2…", "name": "New chat" }
```
`name` optional. Conversation IDs are server-generated 32-hex opaque strings.
### `conversation.open` — load history (reconnect recovery)
```json
→ { "type": "conversation.open", "conversation_id": "a1b2…", "limit": 200, "request_id": "r3" }
← { "type": "conversation.history", "conversation_id": "a1b2…", "request_id": "r3",
"messages": [
{ "message_id": "m12", "role": "user", "text": "hey", "ts": "…|null" },
{ "message_id": "m13", "role": "assistant", "text": "hi!", "ts": "…|null" } ],
"has_more": true,
"attachments": [],
"run": null,
"tools": [],
"approvals": [],
"clarifications": [] }
```
History is the authoritative Hermes transcript (`role` is always `user` or
`assistant`). It spans all durable reset and compression generations sharing
the Pheby conversation key. `limit` is capped at 500; the response returns the
newest page in chronological order. While `has_more` is true, request earlier
pages with `"before_message_id": "m12"`, using the first message ID from the
previous page as an exclusive cursor. This is display history only; model
context can still be compressed. The other fields form a recoverable snapshot: unexpired
attachments, the active run (if any), latest structured tool states, and
pending approval/clarification requests. On reconnect, replace local state
with this snapshot, then consume new live events. Unknown conversation →
`conversation_not_found` error.
### `conversation.rename`
```json
→ { "type": "conversation.rename", "conversation_id": "a1b2…", "name": "Renamed", "request_id": "r4" }
← { "type": "conversation.renamed", "conversation_id": "a1b2…", "name": "Renamed", "request_id": "r4" }
← { "type": "conversation.renamed", "conversation_id": "a1b2…", "name": "Renamed" } // broadcast
```
### `conversation.delete`
```json
→ { "type": "conversation.delete", "conversation_id": "a1b2…", "request_id": "r5" }
← { "type": "conversation.deleted", "conversation_id": "a1b2…", "request_id": "r5" }
```
Deletes the Hermes session transcript and the routing entry. Attachments
belonging to the conversation age out on their own 7-day schedule.
**Not supported by design:** message editing, per-message deletion,
regeneration, edit-and-resend. If you need to "undo", send a correction
message (the agent sees the whole transcript).
---
## Chat & streaming
### `chat.send`
```json
→ { "type": "chat.send", "conversation_id": "a1b2…", "text": "What's the weather?", "request_id": "r6" }
```
`attachment_ids` is an optional list of up to 10 distinct inbound upload IDs
from `POST /attachments` in the **same conversation**. The adapter rejects
unknown, already-sent, or cross-conversation IDs. Upload first, then include
all IDs in the single `chat.send`; a rejected active run leaves them retryable.
The `text` field is still required and nonblank (for file-only sends, provide
a short caption). A successful send anchors the attachments to the user turn
and broadcasts `attachment.added`. Small text files are included in agent
context; other files remain available to the agent as local media paths.
### Server → Client run lifecycle
```json
← { "type": "run.accepted", "conversation_id": "a1b2…", "run_id": "8c1f…", "request_id": "r6" }
← { "type": "message.start", "conversation_id": "a1b2…", "run_id": "8c1f…", "message_id": "draft-8c1f…" }
```
While the agent streams, the server pushes **cumulative** draft text (the
client can simply replace the bubble's text each time — no delta stitching):
```json
← { "type": "message.delta", "conversation_id": "a1b2…", "message_id": "draft-8c1f…",
"text": "It's currently 27°C…", "ts": "..." }
```
Completion (final text supersedes the draft — render the final, drop the
draft):
```json
← { "type": "message.complete", "conversation_id": "a1b2…",
"message_id": "draft-8c1f…", "text": "…full final answer…", "ts": "..." }
```
`message.complete` with `"kind": "notice"` is a gateway lifecycle/status
notice rather than conversation content — render or ignore.
Run end:
```json
← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…",
"status": "completed" | "cancelled" | "failed",
"error": "only on failure" }
```
State machine per assistant turn:
`run.accepted → message.start → (message.delta)* → message.complete → run.finished`.
A non-streaming turn may skip `message.delta`; `message.start` is always sent
when the run is accepted. Only one run may be active per conversation;
another `chat.send` receives `run_active`. Never infer state from text.
### `run.cancel` — stop an active run
```json
→ { "type": "run.cancel", "conversation_id": "a1b2…", "run_id": "8c1f…", "request_id": "r7" }
← { "type": "run.finished", "conversation_id": "a1b2…", "run_id": "8c1f…", "status": "cancelled" }
```
Cancellation uses Hermes's supported interrupt mechanism (agent interrupt +
run-generation invalidation) — the conversation stays consistent and
resumable. A stale `run_id`, missing run, or run that Hermes can no longer
interrupt returns `not_found`; success emits exactly one cancelled event.
---
## Tool events (structured — never chat text)
Tool activity arrives as `tool.event` messages, completely separate from
`message.*` chat content. The client renders them as compact tool components
(ChatGPT-style) attached to the assistant turn.
```json
{ "type": "tool.event",
"conversation_id": "a1b2…",
"run_id": "8c1f…",
"tool_call_id": "t-1a2b3c4d5e6f",
"tool_name": "web_search",
"status": "running" | "completed" | "failed" | "cancelled",
"args_redacted": { "query": "cats" }, // only on "running"; secret-looking keys redacted
"duration_ms": 1234, // only on completion/failure (may be null)
"error": "only on failed, truncated", // only on "failed"
"ts": "..." }
```
Correlate `running` → `completed`/`failed` by `tool_call_id`. Both events use
Hermes's authoritative call ID from the pre/post tool hooks and include the
active `run_id` when one exists. Secret-looking argument keys are redacted
recursively before leaving the server.
No fake "Searching the web…" text is ever injected into `message.*` events.
---
## Approvals
When Hermes pauses for a human decision on a dangerous action:
```json
{ "type": "approval.request",
"approval_id": "3d4e5f6070a1",
"session_key": "agent:main:pheby:dm:a1b2…",
"conversation_id": "a1b2…",
"run_id": "8c1f…",
"command": "rm -rf /tmp/build-output",
"description": "Destructive shell command (rm -rf)",
"choices": ["once", "session", "always", "deny"],
"ts": "..." }
```
Respond:
```json
→ { "type": "approval.respond", "approval_id": "3d4e5f6070a1",
"choice": "once" | "session" | "always" | "deny",
"reason": "optional free text with deny", "request_id": "r8" }
← { "type": "approval.resolved", "approval_id": "3d4e5f6070a1",
"choice": "once", "accepted": true, "request_id": "r8" }
```
The single resolution event is broadcast to every connected client; the
responding client's `request_id` is included on that broadcast.
Choices map to Hermes semantics: `once` (approve this action), `session`
(approve pattern for this conversation), `always` (also persist), `deny`
(decline; the agent is told NOT to retry). Unknown/stale ID →
`approval_not_found` error. Hermes itself fails the approval closed after its
own timeout, so a silently-closed socket can't leave a zombie gate.
---
## Clarifications / choices
```json
{ "type": "clarify.request",
"clarify_id": "c1a2b3d4e5",
"session_key": "agent:main:pheby:dm:a1b2…",
"conversation_id": "a1b2…",
"run_id": "8c1f…",
"question": "Deploy to staging or production?",
"choices": ["staging", "production"], // null ⇒ free text only
"allow_free_text": true,
"ts": "..." }
```
Respond (either a choice value or free text):
```json
→ { "type": "clarify.respond", "clarify_id": "c1a2b3d4e5", "response": "production", "request_id": "r9" }
← { "type": "clarify.resolved", "clarify_id": "c1a2b3d4e5",
"accepted": true, "request_id": "r9" }
```
The single resolution event is broadcast. A stale/timed-out ID instead gets
`clarify_not_found`. Always render an "Other" affordance — Hermes
clarifications accept free text.
---
## Auto-approve (YOLO)
YOLO is always conversation-scoped. It skips ordinary dangerous-command
approval prompts for that Hermes session, while Hermes hard deny rules and
hardline safety floors still apply.
```json
→ { "type": "yolo.current", "conversation_id": "a1b2…", "request_id": "r15" }
← { "type": "yolo.snapshot", "enabled": false, "scope": "conversation",
"conversation_id": "a1b2…", "request_id": "r15" }
→ { "type": "yolo.set", "conversation_id": "a1b2…", "enabled": true,
"request_id": "r16" }
← { "type": "yolo.changed", "enabled": true, "scope": "conversation",
"conversation_id": "a1b2…", "request_id": "r16" }
// plus broadcast of yolo.changed (without request_id) to all clients
```
`enabled` must be a JSON boolean. Unknown conversations return
`conversation_not_found`. The adapter also persists the flag in Hermes session
metadata so it can be restored after the gateway restarts.
---
## Attachments (both directions)
### Client → agent uploads
```
POST /attachments?conversation_id=a1b2…&filename=notes.md
Authorization: Bearer <PHEBY_SECRET>
Content-Type: text/markdown
# Raw file bytes, not multipart or JSON
```
Returns `201 {"attachment": {"attachment_id": "<32 hex chars>", ...}}`.
The descriptor contains `direction: "inbound"`, `message_id: null`, and the
same fields as agent deliverables below. An upload is **not** broadcast or
listed in conversation history until `chat.send` successfully claims it;
an unused upload expires with normal retention. Limit: 64 MiB per file.
Missing credentials → 401; bad conversation ID or empty body → 400;
oversize body → 413. The file picker may select up to 10 files per message.
The server keeps only registered copies in adapter-owned storage and never
trusts a client-provided path. Markdown and other small `text/*` files
(up to 100 KiB) are inlined into the agent's turn; the conversation history
returns just the user's original message text.
### Agent → client deliverables
When the agent produces a file (image, document, audio, video… via Hermes's
normal `MEDIA:` deliverable pipeline), the server copies it into
adapter-managed storage and broadcasts:
```json
{ "type": "attachment.added",
"conversation_id": "a1b2…",
"attachment": {
"attachment_id": "e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0",
"filename": "report.pdf",
"mime_type": "application/pdf",
"size": 48213,
"kind": "image" | "voice" | "video" | "audio" | "document",
"inline_image": false,
"direction": "outbound",
"conversation_id": "a1b2…",
"message_id": "draft-8c1f…" | null,
"created_at": "2026-09-02T18:30:00+00:00",
"expires_at": "2026-09-09T18:30:00+00:00", // null when retention=0
"download_path": "/attachments/e5f6…" },
"ts": "..." }
```
Download over **HTTPS** (authenticated — same `PHEBY_SECRET`):
```
GET {download_path}
Authorization: Bearer <PHEBY_SECRET> (or ApiKey <secret>, or X-Pheby-Secret: <secret>)
```
* `inline_image: true` → `kind == "image"`, safe for an inline preview
(`BitmapFactory` / `AsyncImage` with the same authenticated GET).
* Everything else: download and open as an Android document.
* Unknown / expired / malformed ID → `404` with
`{"error":{"code":"not_found","message":"Attachment unavailable"}}` —
no implementation details.
* **The client can never request arbitrary files** — only registered
attachment IDs resolve.
* **Retention:** adapter copies expire after 7 days (configurable) and are
deleted by an hourly cleanup. Original files the agent produced elsewhere
on the host are never touched. `expires_at` tells the client when to stop
offering the download.
---
## Models
### List providers + models
```json
→ { "type": "models.list", "conversation_id": "a1b2…", "request_id": "r10" }
← { "type": "models.snapshot", "request_id": "r10",
"providers": [
{ "slug": "openrouter", "name": "OpenRouter", "is_current": true,
"models": ["z-ai/glm-5.3-flash", "anthropic/claude-sonnet-4", "…"],
"total_models": 42 } ],
"current_model": "z-ai/glm-5.3-flash",
"current_provider": "openrouter",
"supported_reasoning_efforts": ["minimal","low","medium","high","xhigh","max","ultra"],
"ts": "..." }
```
Lists come from Hermes's own credential-aware picker data — nothing is
hardcoded. Models are exactly what the configured providers expose. The
optional `conversation_id` makes `current_model`, `current_provider`, and
`scope` reflect that conversation's override.
### Read / change current model
```json
→ { "type": "models.current", "conversation_id": "a1b2…", "request_id": "r11" }
← { "type": "model.current", "request_id": "r11", "model": "z-ai/glm-5.3-flash",
"provider": "openrouter", "scope": "conversation",
"conversation_id": "a1b2…", "ts": "..." }
→ { "type": "model.set", "model": "anthropic/claude-sonnet-4",
"provider": "anthropic", // optional
"conversation_id": "a1b2…", // present ⇒ session-scoped override
"request_id": "r12" }
← { "type": "model.changed", "model": "anthropic/claude-sonnet-4",
"provider": "anthropic", "scope": "conversation" | "global", "request_id": "r12" }
// plus broadcast of model.changed (without request_id) to all clients
```
Omit `conversation_id` ⇒ the change is persisted globally (Hermes
`model.default`). A `model.changed` with `scope:"global"` tells every open
conversation the default moved.
---
## Reasoning effort
```json
→ { "type": "reasoning.current", "conversation_id": "a1b2…", "request_id": "r13" }
← { "type": "reasoning.snapshot", "request_id": "r13",
"effort": "medium", // current effective effort (may be null = provider default)
"enabled": true, // false ⇒ thinking disabled
"scope": "conversation", "conversation_id": "a1b2…",
"supported_efforts": ["none","minimal","low","medium","high","xhigh","max","ultra"],
"ts": "..." }
→ { "type": "reasoning.set", "effort": "high", "conversation_id": "a1b2…", "request_id": "r14" }
← { "type": "reasoning.changed", "effort": "high", "scope": "conversation", "request_id": "r14" }
```
`effort: "none"` disables thinking. Invalid values → `bad_request`. Scope
rules mirror `model.set` (with `conversation_id` ⇒ session override; without ⇒
global `agent.reasoning_effort`). **Capability note:** Hermes knows *whether*
a model supports reasoning (models.dev metadata) but does not expose a
per-provider enum of valid effort values; the listed levels are Hermes's
canonical set — unsupported levels on a given provider surface as a provider
error on the next turn, not at set time. This is a documented Hermes
limitation, not a Pheby guess.
---
## Unsolicited messages & reconnect behavior
The WebSocket stays connected; any Hermes-originated output destined for the
Pheby platform (scheduled/cron deliveries, background completions,
notifications) is pushed as normal `message.*` / `run.*` events even when it
is not a reply to your last request. Reconnection procedure for clients:
1. Reconnect WS, redo `hello`.
2. Re-`conversation.open` the conversations you show; replace local history,
attachments, run/tool state, and pending decisions with its authoritative
snapshot.
3. Re-`models.current` / `reasoning.current` if those views are visible.
4. Live events continue from there.
No external push service exists (no FCM); Android notification behavior is
the client's responsibility while the socket is down.