Skip to content

AbstractGateway — API overview

The HTTP API is implemented with FastAPI under the /api prefix: - Health: GET /api/health (unauthenticated; status, runner, watchdog: {enabled, limit_s, last_tick_age_s} — see troubleshooting.md). The previous process's last watchdog incident is admin-only: GET /api/gateway/host/runner → last_hang: {at, blocked_s, reason, top_frame, dump_path, file, line} or null - Gateway surface: /api/gateway/* (durable runs + operator tooling)

The API is documented at runtime: - OpenAPI JSON: GET /openapi.json - Swagger UI: GET /docs (use Authorize to paste the bearer token)

Context: - In the AbstractFramework ecosystem, UIs and automations call this API to operate AbstractRuntime runs. - Architecture diagram and core concepts: architecture.md

Route families

This page covers the run contract, artifacts, discovery, media, models and host state. Other route families are documented next to the feature they serve:

Routes Purpose Reference
/api/gateway/session/login, /session/logout, /session/claim, /me browser sessions, one-time sign-in links, the current principal security.md, first-run.md
/api/gateway/admin/users, /admin/runtime-reservations user accounts and retained runtimes (admin) security.md
/api/gateway/admin/accounts, /admin/accounts/{id}/active, /admin/accounts/{id}/activity, /me/accounts, /me/accounts/{id}/activity, /me/activity the Accounts page: users and entities in one list (admin), your own account and your entities (everyone), the Active switch, activity from the audit log below
/api/gateway/admin/runtime-config runtime settings (admin) configuration.md
/api/gateway/accounts/{me\|account}/preferences one account's client preferences: the default workflow per app (the account itself, an admin, an entity's creator) below
/api/gateway/workspace/policy, /workspace/policy/{account}, /sessions/{id}/workspaces, /workspace/effective/{account} (GET and dry-run POST), POST /workspace/path-check workspaces, three levels: the gateway policy = the eligible set (read: everyone; write: admin), one account's default subset (admin, the account itself, an entity's creator; me = the caller), one conversation's subset (its owner, or an admin), the effective set the gateway enforces, and the path check run before saving a row ({path} → {path, normalized, absolute, exists, is_dir, valid, sentence}; any signed-in principal) below, security.md
/api/gateway/admin/runtimes, ?account=<id>[&tenant_id=<t>] the runtime inventory (admin); account keeps only the planes that account owns and echoes filter console.md
/api/gateway/network, /network/restart network exposure, addresses, reverse proxy configuration.md
/api/gateway/apps/*, /apps/handover/{code}, /apps/tui-handover, /api/gateway/apps/desktop-handover browser apps, terminal apps, the Assistant and its sign-in apps.md
/api/gateway/runs/{run_id}/workspace, /workspace/files, /workspace/content browse and preview a run's folder (the run's owner) below
/api/gateway/skills, /admin/skills/reseed the skills shelf below, configuration.md
/api/gateway/about versions this gateway runs (no sign-in) below
/api/gateway/engines/* local engine installs engines.md
/api/gateway/models/download*, /models/downloads* model download jobs and their event stream model-downloads.md
/api/gateway/host/* host state, pause, restart, update, tray Host state, Host control
/api/gateway/backlog/*, /reports/*, /triage/*, /processes operator tooling maintenance.md
/api/gateway/entities/* summoned entities entities.md
/api/gateway/automations*, /api/gateway/trigger-sources automations: recurring workflows, their occurrences and attention automations.md

OpenAI API

The OpenAI-compatible API is served at /v1 on this Gateway listener (/v1/models, /v1/chat/completions, /v1/embeddings, ...); callers use their gateway token as the API key. /core/v1 answers 308 to /v1 (deprecated). Status and logs: GET /api/gateway/openai-api[/logs]; admin changes: /api/gateway/admin/core-endpoint. See openai-api.md for the supported surface, access settings and errors.

Auth

By default, /api/gateway/* is protected by GatewaySecurityMiddleware (bearer token + origin allowlist).
See: security.md.

All examples below assume:

export BASE_URL="http://127.0.0.1:8080"
export AUTH="Authorization: Bearer $(cat "$ABSTRACTGATEWAY_DATA_DIR/auth/bootstrap-admin-token")"

Provider connections

Gateway-owned provider connections let users create reusable cloud, local, or OpenAI-compatible endpoints without putting raw API keys in workflow JSON or browser storage. The API route is named provider-endpoint-profiles; the console presents them as provider connections.

  • GET /api/gateway/config/provider-endpoint-profiles: list visible profiles.
  • POST /api/gateway/config/provider-endpoint-profiles: create a user- or admin-owned profile.
  • POST /api/gateway/config/provider-endpoint-profiles/discover-models: discover models for a draft or saved profile by calling the configured provider family and base URL with the entered or server-side key. The raw key is never returned.
  • PUT or DELETE /api/gateway/config/provider-endpoint-profiles/{profile_id}: update or delete a profile.

Enabled profiles appear in GET /api/gateway/discovery/providers as virtual providers such as endpoint:office-vllm. Model discovery through GET /api/gateway/discovery/providers/{provider_name}/models returns either the fixed profile allowlist or the live endpoint model catalog.

Core workflow lifecycle

1) List bundles (bundle mode)

curl -sS -H "$AUTH" "$BASE_URL/api/gateway/bundles"

Each item also says where the bundle came from and what it does:

  • source: shipped (a file the gateway package ships), published (published from AbstractFlow through this gateway: the publish route's stamp in the manifest metadata) or imported (anything else: an uploaded .flow, or a file copied into the folder);
  • description: the default entrypoint's description ("" when it has none).

Upload a bundle:

curl -sS -H "$AUTH" \
  -F "file=@./my-bundle@0.1.0.flow" \
  -F "overwrite=false" \
  -F "reload=true" \
  "$BASE_URL/api/gateway/bundles/upload"

2) Start a run

curl -sS -H "$AUTH" -H "Content-Type: application/json" \
  -d '{"bundle_id":"my-bundle","input_data":{"prompt":"Hello"}}' \
  "$BASE_URL/api/gateway/runs/start"

If you need a specific entrypoint:

curl -sS -H "$AUTH" -H "Content-Type: application/json" \
  -d '{"bundle_id":"my-bundle","flow_id":"ac-echo","input_data":{"prompt":"Hello"}}' \
  "$BASE_URL/api/gateway/runs/start"

Every start answers {run_id, runner_warning, resolved_workflow}. resolved_workflow names the workflow the run really runs: {workflow_id, bundle_id, bundle_version, flow_id, registry_scope, name, source: "gateway_default" | "client", interface}. The same object is kept in the run's inputs as input_data.workflow_selection (written by the gateway; a value sent by the client is replaced), so GET /runs/{run_id}/input_data tells a restored conversation how its workflow was chosen.

Evidence: request/response models live in src/abstractgateway/routes/gateway.py (StartRunRequest, start_run).

The gateway default agent workflow (flow_id: "@default")

An agent client (AbstractCode, the Assistant, the Telegram bridge) can let the gateway choose the workflow for an agent interface:

curl -sS -H "$AUTH" -H "Content-Type: application/json" \
  -d '{"flow_id":"@default","interface":"abstractcode.agent.v1","input_data":{"prompt":"Hello"}}' \
  "$BASE_URL/api/gateway/runs/start"
  • interface is required with @default (400 without it); bundle_id and bundle_version are not accepted with it.
  • The default is resolved at every start, so a change applies to the next new turn of any conversation that sends @default.
  • When the default cannot run (its workflow is gone, deprecated, or does not declare the interface), the start is refused with 409 and a message naming the setting (agents.default_workflow.<interface>) and where its value comes from. The gateway never quietly runs another workflow instead.
  • POST /runs/schedule accepts the same flow_id: "@default" + interface (the schedule then targets the version resolved at that moment).

What the default is for each interface (readable without admin rights): GET /bundles and GET /workflow-catalog carry

"default_agent_workflows": {
  "abstractcode.agent.v1": {"workflow_id": "basic-agent@0.0.5:81795ea9", "bundle_id": "basic-agent",
                            "bundle_version": "0.0.5", "flow_id": "81795ea9", "registry_scope": "private",
                            "name": "basic-agent", "source": "default"}
},
"default_agent_workflows_unavailable": {
  "abstractassistant.agent.v1": {"source": "default", "value": null,
    "reason": "no host workflow declares abstractassistant.agent.v1; the Assistant uses its built-in orchestrator"}
}

and every entrypoint row carries is_agent_default and agent_default_interfaces. These are different from default_bundle_id (the bundle a bare flow_id falls back to), from a catalog record's is_default (its default version) and from a bundle's default_entrypoint.

The setting itself is agents.default_workflow.<interface> (see Configuration). In GET /admin/runtime-config, each agents.default_workflow.<interface> row also carries the plain words the consoles show, from one table in agent_defaults.py (INTERFACE_TABLE):

Field Meaning
interface, label, app, help the interface id, its plain name ("AbstractCode — chat agent"), the app that asks for it (null when none does) and one sentence of help
group apps (an app asks the gateway for it as its default agent) or other (declared by a workflow; no app asks for it by default)
state builtin (nothing saved; the gateway default runs: the shipped bundle, else the newest available workflow), none (nothing saved and no workflow declares the interface; reason says so), set (a saved value that runs) or broken (a saved value that no longer resolves)
value the saved value, else the gateway default, else null
reason for broken: "Broken: workflow bundle 'coding-agent' is not on this gateway — pick another workflow or the gateway default."; for none: "no workflow on this gateway declares "

For VisualFlow bundles, Gateway runs the packed JSON through AbstractRuntime. Structured LLM/Agent schemas are Runtime/Core-owned: response remains textual, and schema-conformant object values are available through the node data output for data edges such as Break Object and Switch.

Durable session replay (use_session_history)

Thin clients do not need to carry conversation transcripts. Passing "input_data": {"use_session_history": true} together with a session_id makes the gateway seed the run's context.messages from the session's prior COMPLETED root runs before the run starts: the run store is the durable transcript.

curl -sS -H "$AUTH" -H "Content-Type: application/json" \
  -d '{"bundle_id":"my-bundle","session_id":"sess-1","input_data":{"prompt":"and what did I say before?","use_session_history":true}}' \
  "$BASE_URL/api/gateway/runs/start"

Rules (the model-vs-display divergence contract — what the model replays is deliberately narrower than what history views display):

  • Client-provided non-empty context.messages always win; the seed never overwrites them. An EMPTY client context.messages list does not count as a transcript — the seed still runs (leave out use_session_history to start without history).
  • Only COMPLETED root runs of the session contribute, as strictly alternating user/assistant pairs. FAILED and CANCELLED turns are invisible to replay by design (a promptless answer or answerless prompt would seed a dangling message and invite re-answering a stale ask); history views still show them.
  • Steering/operator guidance injected mid-run is not replayed.
  • The history window: replay keeps the most recent turns that fit 50,000 estimated tokens (AbstractRuntime HISTORY_REPLAY_MAX_TOKENS), as whole messages, newest first. No message is cut, and there is no message-count or character cap. The model can use the rest of its context window. When older turns are dropped, the oldest replayed message starts with a labeled [#TRUNCATION: ...] line. A newest turn that alone exceeds 50,000 tokens is kept whole (oversize_turn_kept: true).
  • The retired caps input_data.session_history_max_messages and input_data.session_history_max_chars are ignored and listed in _runtime.session_history.ignored_inputs. The environment variables ABSTRACTGATEWAY_SESSION_HISTORY_MAX_MESSAGES and ABSTRACTGATEWAY_SESSION_HISTORY_MAX_CHARS are not read.
  • Failures degrade to a labeled _runtime.session_history #FALLBACK note and an unseeded start — never a blocked run. Success records the window on the run as _runtime.session_history: {seeded, policy, max_tokens, token_estimator, replayed_messages, replayed_tokens, dropped_messages, dropped_tokens, dropped_counts_complete, oversize_turn_kept}. dropped_counts_complete: false means replay stopped reading once the window was full: older turns were dropped too, and they are not counted.
  • Entity lanes never ride this: their transcript authority is the entity home (_visit.history / the chat driver), not the run store.

Evidence: _seed_session_history in src/abstractgateway/hosts/bundle_host.py and abstractruntime.session_history.session_chat_messages.

2b) Schedule a run (bundle mode)

POST /api/gateway/runs/schedule starts a scheduled parent run that launches the target workflow as child runs over time.

Example (run 3 times, every hour, starting now):

curl -sS -H "$AUTH" -H "Content-Type: application/json" \
  -d '{"bundle_id":"my-bundle","flow_id":"ac-echo","input_data":{"prompt":"Ping"},"start_at":"now","interval":"1h","repeat_count":3,"share_context":true,"session_id":"sess-1"}' \
  "$BASE_URL/api/gateway/runs/schedule"

Notes: - start_at: ISO 8601 timestamp (recommended) or "now". - interval: e.g. "15m", "1h", "2d". If omitted, runs once. - repeat_count: if omitted and interval is set, repeats forever. Alternatively use repeat_until (ISO 8601). - To stop a schedule, cancel the scheduled parent run via POST /api/gateway/commands with type cancel.

Evidence: ScheduleRunRequest, start_scheduled_run in src/abstractgateway/routes/gateway.py.

2c) Shared workflow catalog

Private /api/gateway/bundles routes are scoped to the signed-in user's routed runtime, and you may change the registry you own. The gateway's own bundle directory is shared by every user, so writing it — upload, delete, reload, deprecate, and POST /visualflows/{flow_id}/publish — requires an admin principal and otherwise returns 403. Listing and running are unaffected. See security.md for the full rule.

Shared/default workflows use the Gateway workflow catalog instead:

curl -sS -H "$AUTH" "$BASE_URL/api/gateway/workflow-catalog"

Admin-only catalog operations live under /api/gateway/admin/workflow-catalog/*:

  • upload or promote immutable .flow versions;
  • move a bundle's default pointer;
  • set ACLs;
  • deprecate, block, or tombstone a version without deleting bundle bytes.

Start a catalog workflow in the requesting user's runtime by setting registry_scope:

curl -sS -H "$AUTH" -H "Content-Type: application/json" \
  -d '{"registry_scope":"tenant_catalog","bundle_id":"basic-agent","flow_id":"root","input_data":{"prompt":"Hello"}}' \
  "$BASE_URL/api/gateway/runs/start"

If bundle_version is omitted, Gateway uses the admin-managed catalog default pointer. Exact older versions keep working until that specific version is deprecated, blocked, or tombstoned.

Catalog scope is explicit: omitting registry_scope starts only private runtime bundles. Flow/schema inspection for catalog workflows should use the ACL-aware catalog endpoints:

  • GET /api/gateway/workflow-catalog/{bundle_id}/versions/{bundle_version}/flows/{flow_id}
  • GET /api/gateway/workflow-catalog/{bundle_id}/versions/{bundle_version}/flows/{flow_id}/input_schema

framework_catalog is reserved but not loadable yet; use tenant_catalog.

2d) Docs Q&A (docs-qa catalog bundle)

docs-qa is the shared transport for docs-grounded assistant panels (the unified top-bar drawers). The contract: the CALLER supplies its own corpus (typically its llms.txt text) — the bundle never guesses one, so answers are never silently grounded on another app's docs.

Fresh installs need no manual publish: the gateway ships docs-qa in the wheel and boot idempotently publishes it into the tenant catalog (publish-if-absent by exact version; an admin's default pointer, tombstones, and publisher attribution are never touched; publisher system:gateway-boot). The publish is skipped on custom-bundle deployments whose private registry carries no LLM-bearing flow (it would add a boot requirement they never had) and can be disabled with ABSTRACTGATEWAY_AUTO_PUBLISH_SHIPPED=0 — the manual upload below then remains the path.

curl -sS -H "$AUTH" -H "Content-Type: application/json" -d '{
  "registry_scope": "tenant_catalog",
  "bundle_id": "docs-qa",
  "bundle_version": "0.1.1",
  "flow_id": "docsqa001",
  "session_id": "myapp-docs-assistant:<one id per conversation>",
  "input_data": {
    "prompt": "How do I publish a workflow bundle?",
    "docs": "<your llms.txt text>",
    "app": "MyApp",
    "use_session_history": true
  }
}' "$BASE_URL/api/gateway/runs/start"

Conversation history comes from the run's session, never from the caller (0.1.1): start every question of one conversation with the same session_id and use_session_history: true. The gateway replays the session's earlier turns through the runtime's history window (the newest whole turns up to 50,000 tokens) and the bundle's LLM call includes them; a new conversation is a new session_id. 0.1.0's question + history inputs (last 12 messages kept) are gone.

Then poll GET /runs/{run_id} (or stream the ledger); the answer is output.response, and session_history is the window's receipt (replayed_messages, dropped_messages, ...) to show when earlier messages were not replayed. provider/model/temperature may ride input_data to override gateway defaults. Answers cite section headings and say plainly when the docs do not answer — the bundle refuses to invent endpoints or behavior. Docs Q&A must never route through entity chat (a visit is billable and forms memories).

Run-level skills selection

input_data.skills (a list of skill NAMES) attaches curated skills to any run started through /runs/start:

curl -sS -H "$AUTH" -H "Content-Type: application/json" -d '{
  "bundle_id": "basic-agent",
  "input_data": {"prompt": "…", "skills": ["agora-collaboration"]}
}' "$BASE_URL/api/gateway/runs/start"

Trust semantics (the same abstractskill gate as GET /skills and the workforce spawn lane — one gate, never a second resolver): VALIDATED skills activate and their index lands in the run's _runtime.skills_block (byte-stable for the whole run) with the read_skill tool made reachable; UNVERIFIED skills are held; advisory-BLOCKED skills never ride. Every outcome is recorded as a labeled verdict in _runtime.skills_resolution (requested/active/verdicts/resolved_tree_hashes) — nothing is silently dropped. Agent-node subruns inherit the block verbatim with read_skill appended to explicit child allowlists (empty allowlists keep registry defaults). A caller-supplied _runtime.skills_block is never overwritten; the selection is then ignored with a labeled verdict.

The gateway serves the corpus every Docs assistant grounds on (the kit's DocsAssistantDrawer in the console and the apps):

curl -sS -H "$AUTH" "$BASE_URL/api/gateway/docs/corpus"            # the gateway's own llms.txt
curl -sS -H "$AUTH" "$BASE_URL/api/gateway/docs/corpus?app=code"   # an app's llms.txt

Returns {app, source, chars, text}. With app=<id> (code, flow, observer, continuum, entity) the text is the llms.txt the running app serves from its own build (GET /llms.txt on its loopback port, text/plain only; source app:<id>:llms.txt); an app that is not running, an unknown id, or an app that does not serve one is a 404 that says which. Without app (or app=gateway): Resolution order: the ABSTRACTGATEWAY_DOCS_CORPUS env override first (set-but-missing is an honest 404 naming the checked candidates, never a silent fallback), then the repo llms.txt in dev checkouts, then the corpus packaged with the wheel.

Skills shelf

GET /api/gateway/skills lists the skills of the gateway's shelf: {skills: [...], shelf, shelf_source, bundled_version, warnings}. shelf_source says where the shelf comes from: stored (the saved setting skills.shelf), env (a legacy launch environment value), seeded (the gateway's own copy in <data dir>/skills/registry, kept up to date from the curated shelf that ships with AbstractSkill at each start), checkout (a framework checkout, used only when the gateway's own copy is missing) or none. An empty list always comes with a warning that says why and what to do; warnings are plain sentences meant to be shown as they are.

POST /api/gateway/admin/skills/reseed (admin) refreshes the gateway's own copy now and answers the seed report (added, updated, unchanged, the kept_* lists with what was kept and why, bundled_version, previous_version). An edit made in that folder is never overwritten.

3) Replay the ledger (cursor-based)

Ledger pages are replayed using after as “number of items already consumed”.

curl -sS -H "$AUTH" "$BASE_URL/api/gateway/runs/<run_id>/ledger?after=0&limit=200"

Response shape: - items: list of durable ledger records - next_after: the next cursor to use

Evidence: src/abstractgateway/routes/gateway.py (get_ledger).

3b) Replay ledgers for multiple runs (batch)

Use POST /api/gateway/runs/ledger/batch to reduce request fanout when observing many runs/subflows.

curl -sS -H "$AUTH" -H "Content-Type: application/json" \
  -d '{"limit":200,"runs":[{"run_id":"<run_id_1>","after":0},{"run_id":"<run_id_2>","after":0}]}' \
  "$BASE_URL/api/gateway/runs/ledger/batch"

Evidence: src/abstractgateway/routes/gateway.py (get_ledger_batch).

4) Stream ledger updates (SSE)

SSE is an optimization; clients should always be able to reconnect by replaying from the last next_after.

curl -N -H "$AUTH" "$BASE_URL/api/gateway/runs/<run_id>/ledger/stream?after=0"

Each ledger record arrives as event: step with an id: line (the records consumed so far); a reconnect sends Last-Event-ID (or ?after=) and resumes there. event: done closes the stream once the run is finished and everything was sent.

Evidence: src/abstractgateway/routes/gateway.py (stream_ledger).

4b) Live replies: token deltas on the same stream

A run started with input_data._runtime.stream: true also sends the model's reply while it is being written, on the same SSE stream:

event: llm.delta
data: {"kind":"llm.delta","run_id":"…","root_run_id":"…","parent_run_id":null,"node_id":"llm","call_id":"<step id>","seq":3,"text":" a natural","channel":"content","snapshot":false}

event: llm.delta_end
data: {"kind":"llm.delta_end","run_id":"…","root_run_id":"…","parent_run_id":null,"node_id":"llm","call_id":"<step id>","seq":24,"reason":"completed","snapshot":false}
  • call_id is the step id of the LLM call; the durable llm_call record with the same step_id carries the final answer and replaces the live text.
  • channel is content (the answer) or reasoning (thinking text, already separated on the server; clients do not parse <think> themselves).
  • seq counts the events of one call (deltas and end).
  • Delta frames have NO id: line: they are not ledger records, and Last-Event-ID never moves because of them.
  • Snapshots. When a client connects (or reconnects) while a call is still being written, it first receives one frame per open call and channel with snapshot: true and the text so far, then the live frames. A client drops its live bubbles on every (re)connect and applies the snapshots.
  • Order. A call's llm.delta_end is sent after its durable record; the last short batch of text can arrive just after the record, so a client ignores deltas of a call whose record it already holds. done comes after every live frame.
  • Sub-runs. A stream of a run also carries the deltas of every run below it (a coding agent's delegated calls); a stream of a child run carries that child's subtree only. run_id names the run that wrote the text, root_run_id its root.
  • Endings. reason is completed, failed, cancelled or unavailable. unavailable comes with a detail saying why that call did not stream (structured_output, remote_core, provider_cannot_stream, usage_unavailable, prompt_cache_unavailable, node_stream_off, sink_error); the answer is complete either way. When a run ends (stopped, failed) while a call is still open, the gateway closes that call with reason: "cancelled" or "failed" and synthetic: true.
  • Nothing is capped: a call's text is kept whole until it ends, and a slow client catches up from it.
  • Another user's run answers 404, and no delta of it ever reaches anyone else.

The switch is input_data._runtime.stream (true or false; anything else is refused with 400). When a POST /runs/start request does not say, the gateway setting agents.streaming_default decides (off unless saved); it is never applied to POST /runs/schedule, the bridges or the entity loop. A flow node whose LLM call sets stream: false is never streamed (its end says detail: "node_stream_off"). GET /api/gateway/discovery/capabilities advertises the feature to every client:

"streaming": {"deltas": true, "default": false, "run_field": "_runtime.stream",
              "events": ["llm.delta", "llm.delta_end"], "subtree": true,
              "endpoint": "/api/gateway/runs/{run_id}/ledger/stream",
              "end_reasons": ["completed", "failed", "cancelled", "unavailable"]}

With serve --no-runner plus abstractgateway runner, the runner writes a run's deltas to <data dir>/live/<root run id>.deltas.jsonl (readable by the gateway's user only) and the API process reads them from there; the file is deleted when the run ends (however it ends: completed, failed, stopped, or stopped by the kill switch), and files of finished runs are removed at start. Disk use: the file holds every delta of every call of the run and of its sub-runs, and grows until the ROOT run ends; nothing is capped, so a long agent run that makes many calls (a coding agent working for an hour) can leave a file of several megabytes in <data dir>/live while it runs.

Evidence: src/abstractgateway/live_deltas.py, src/abstractgateway/routes/gateway.py (stream_ledger, _apply_run_stream_switch, discovery_capabilities), src/abstractgateway/hosts/bundle_host.py (install_live_delta_sink).

A run's workspace folder (browse and preview)

A run works in a folder on the gateway computer: the conversation's own folder the gateway made (<data dir>/workspaces/session-…), or the folder the client was started from. A run started without workspace_root gets the conversation's folder (a run without a session gets its own <data dir>/workspaces/<run>), and its file tools resolve relative paths there, never in another workspace. This private workspace is always read & write for its run and is never listed. The agent's workspace context lists the run's allowed workspaces with their paths and modes (security.md). Three routes let the person who started the run see it; another user's run id answers 404.

GET /api/gateway/runs/{run_id}/workspace:

{"run_id": "…", "workspace_root": "/Users/me/Library/Application Support/abstractgateway/workspaces/session-chat-1-3f2a…",
 "kind": "session", "session_id": "chat-1", "exists": true,
 "host": {"hostname": "studio.local", "caller_is_this_machine": true},
 "open_supported": true}

kind is session, run or launch_folder. open_supported is true only for an admin sitting at the gateway computer, the only caller for whom POST /runs/{run_id}/workspace/open can open the folder; elsewhere show the path and the host name.

GET /api/gateway/runs/{run_id}/workspace/files?path=<folder>&recursive=false&limit=<n>:

{"path": "src", "entries": [{"name": "a.py", "path": "src/a.py", "type": "file", "size_bytes": 11, "mtime": "2026-09-25T10:00:00Z"}],
 "truncated": false, "limit": null, "recursive": false,
 "hidden": {"outside_links": 0, "blocked": 0, "other": 0}}

Folders come first, then files, by name. limit is optional; when the listing stops there, truncated is true. hidden counts what is not shown: links that lead outside the folder, and entries the workspace deny list blocks.

GET /api/gateway/runs/{run_id}/workspace/content?path=<file> streams the whole file with its content type, Content-Disposition: inline, X-Content-Type-Options: nosniff and Content-Security-Policy: sandbox, and honours Range (206; 416 outside the file).

Paths are relative to the folder: absolute paths and .. are refused (400), a link that leads outside is refused (403), the gateway's own marker file is never listed nor served. The built-in deny list (credential folders such as ~/.ssh, and the gateway's data folder; see Configuration) is never listed, and reading inside it answers 404, even when the run's folder contains it. The run's effective workspaces are applied again at every call, and nothing else in the gateway's data folder is ever served. A run cannot be started with a workspace_root inside the gateway's data folder either, except the conversation folder the gateway made for the same user.

Workspaces

Three levels, the same shape at each: a posture (allowed_only = "Deny everything, allow listed workspaces", any_except_denied = "Allow everything, refuse listed workspaces"), a default_mode (ro | rw, the mode of every unlisted directory under "Allow everything…") and rows {path, mode: "ro" | "rw" | "deny"} (Read-only, Read & write, Refused). There is no shared workspace. Full model and enforcement: security.md.

Refusals. Every workspace write, the dry run and every run start refuse with HTTP 400 and one shape:

{"detail": {"reason": "workspace_refused", "message": "The gateway allows this workspace read-only: /archive.", "path": "/archive"}}

path names the offending path, or is null. Clients show message (plus "Not saved." after a write). A permission refusal is a 403 whose detail is a sentence.

Gateway: the eligible set (admin)

GET /api/gateway/workspace/policy (any signed-in principal) → {ok, policy} with policy = {posture, default_mode, folders: [{path, mode}], builtin_refused[] (read-only), max_attachment_bytes (read-only), summary}. Each row's mode is the CAP: no level below may reach outside the set or raise a cap. A fresh gateway answers any_except_denied, rw, no rows. The built-in refusals (the gateway's data folder, credential folders) always apply; a non-admin gets builtin_refused: [] and builtin_refused_hidden: true (those paths name the gateway's home and data folder).

PUT /api/gateway/workspace/policy (admin): any subset of {posture, default_mode, folders}. Named fields replace and the others are kept, so one control is one request. Every path is checked like POST /workspace/path-check (absolute, existing). Every row states its mode. These are refused:

  • a duplicate row;
  • a read & write or read-only row inside a built-in refusal (the data folder, except a conversation folder <data>/workspaces/<name> or <data>/users/<tenant>/<runtime>/runtime/workspaces/<name>, and the credential directories): Workspaces: '<path>' is inside the built-in refused workspace '<built-in>'. (the account, session and run levels and the dry run answer the same sentence);
  • an unknown posture or mode;
  • shared_workspace: "shared_workspace no longer exists: list it as a workspace (folders: [{path, mode: "rw"}]). Nothing was saved.";
  • any other field: the earlier allowed_folders, never_allowed, allow_any_folder, launch_folder_trust, mode, client_workspace_scope_overrides, … are all refused by name.

Rows may nest in any combination: the most specific row wins (the longest real-path prefix, refused rows included). A refused /Users/me with an allowed /Users/me/projects (rw) is valid, and so is a refused row inside an allowed one.

The PUT answers like the GET. The GET also carries command_sandbox (round 12), the host's command sandbox state: {state: "sandboxed"|"partial"|"unsandboxed"|"refused", kind, line, sentence, unsandboxed_commands_allowed, configured, flag}. line is shown verbatim (for example "Commands sandboxed: macOS sandbox-exec") and sentence is its tooltip; see security.md.

Account: the account's default subset

GET /api/gateway/workspace/policy/{account} ({account} = name, tenant:name, an entity's slug, or me) → {ok, policy: {account, configured, posture, default_mode, folders}, gateway: <gateway policy>, effective: <effective at the account level>, can_edit}. configured: false means "follow the gateway policy"; the policy block then carries the gateway's posture and default with no rows. 404 for an account that does not exist.

PUT /api/gateway/workspace/policy/{account}: {configured: false} follows the gateway policy again; any of {posture, default_mode, folders} configures the account (unnamed fields keep the stored value, else start from the gateway's posture and default with no rows). A read-only or read & write row must lie inside the eligible set (under "Deny everything…", under a listed row; under "Allow everything…", outside every refused row and built-in refusal) at most at its cap; otherwise the PUT is refused with " is outside the workspaces the gateway allows ()." or "The gateway allows this workspace read-only: .". A Refused row only needs an existing directory: a refusal never widens. Who: an admin for any account; a person for their own; an entity's workspaces by an admin or the entity's creator ("Only an admin or 's creator can change its workspaces." otherwise). The PUT answers like the GET. /workspace/policy/self answers 410 naming /workspace/policy/me.

Session: one conversation's subset

GET /api/gateway/sessions/{session_id}/workspaces → {ok, policy: {session_id, account, configured, posture, default_mode, folders}, gateway, account_default: <effective at the account level>, effective: <effective at the session level>}. It works before the session's first run (nothing stored = configured: false, "Use my default", whose display base is the account default). PUT takes the account level's body and checks: {configured: false} resets; rows are checked against the GATEWAY's eligible set, not the account default (a conversation may use a workspace the account default does not list). The gateway stores the choice in the owner's plane (<plane data dir>/session_workspaces.json), so every app opening the conversation sees it and it survives restarts. Only the session's owner, or an admin with ?account=<tenant:name>, reads or changes it. Audited (workspace_policy_changed, scope session).

Effective set

GET /api/gateway/workspace/effective/{account}[?session=<id>] → {ok, account, session_id, level: "run" | "session" | "account" | "gateway", posture, default_mode, folders: [{path, mode, cap, source: "gateway" | "account" | "session" | "run"}], summary, gateway_summary}. level is the level that applies (session > account > gateway); each mode is the lower of the gateway's cap and the level's rule; default_mode is always ro or rw (under "Deny everything…" it applies to nothing). summary and gateway_summary are the lines every surface shows verbatim:

Deny everything, allow listed workspaces · /Users/me/Pictures (rw) · /Users/me/Documents (ro)
Allow everything, refuse listed workspaces (rw) · /secrets (refused) · /archive (ro)

The default mode appears in parentheses only under "Allow everything…".

POST /api/gateway/workspace/effective/{account} is a dry run: body {workspace: {posture, default_mode, folders} | null, session?: <id>} → the same shape (level: "run" for a payload; null = what a run would get). Nothing is stored; a payload a run start would refuse is refused the same way.

Run start

Every door (POST /runs/start, schedules, host.start_run for bridges and entities, automation occurrences) resolves a one-off workspace > the session's choice > the account default > the gateway policy, clamps it to the eligible set and caps, and sets the run's workspace arguments (workspace_access_mode, allowed, read-only, writable and ignored paths). The one-off is input_data.workspace (POST /runs/start also takes a top-level workspace, moved there). With a session_id it wins for that run and is saved onto the session when the session has no choice yet. At the HTTP doors a row outside the eligible set or above its cap is refused (400 workspace_refused); the in-process doors drop or lower it, never widen, and record each such row on the run as _gateway_workspace.clamped: [{path, asked, got, sentence}] (got null = dropped; the sentence is the HTTP door's). GET /runs/{run_id}/workspace also answers workspace_level, workspace_summary and workspace_clamped. A legacy workspace_allowed_paths list may only narrow; a workspace_root (a launch folder, for example) must be reachable; a client workspace_access_mode: "all_except_ignored" is refused. The run records the level under its vars _gateway_workspace.{level, summary}. An automation keeps its choice in target.input_data.workspace; a definition that sends the older workspace_allowed_paths list gets it converted ("Deny everything, allow listed workspaces", each workspace at its gateway cap). An automation saved with Use my default (no payload) stores workspace: {"configured": false} and follows its owner's default: each occurrence resolves session > account > gateway when it is admitted (its run start), so a wider default applies from the next run, with no revision. Nothing of the default is stored in the definition (only a fail-closed fallback: the built-in denies and workspace_access_mode: "workspace_only"); derived keys a client echoes back with configured: false (an old snapshot) are ignored, never a narrowing. Every automation's workspaces, an explicit choice included, are guarded again at each occurrence against the policy as it is then: a workspace the admin refused or capped since is dropped or lowered for that run and recorded with its sentence (_gateway_workspace.clamped); the keys derived when the definition was saved are never used.

Artifacts and filesystem handoff

Gateway artifacts are the cross-package representation for files, media, and large payloads. Thin clients should pass artifact refs across runs instead of raw bytes or local paths:

{
  "$artifact": "abc123",
  "artifact_id": "abc123",
  "run_id": "session_memory_sess-1",
  "content_type": "image/png",
  "filename": "input.png"
}

Gateway uses three distinct file-like source terms:

  • Artifact: a durable runtime-owned payload reference.
  • Local File: a browser/client upload source. Hosted clients should upload bytes; browser-local paths are never interpreted as server paths.
  • Server File / Server Folder: user-facing wording for a workspace-scoped server path under Gateway policy. The engineering contract is the canonical WorkspacePath string returned by /files/*, artifact import/export, and Runtime file nodes.

Hosted local uploads stay artifact-backed:

  • one local file upload creates one artifact ref;
  • multiple local files create an ordered list of artifact refs in Flow;
  • a local folder uploads one artifact per file and may send source_path (for example reports/2026/summary.md) so relative member paths survive in artifact provenance without exposing browser-local absolute paths.

Upload a local file or folder member:

curl -sS -H "$AUTH" \
  -F "session_id=sess-1" \
  -F "source_path=reports/summary.md" \
  -F "file=@./summary.md" \
  "$BASE_URL/api/gateway/attachments/upload"

List run artifacts:

curl -sS -H "$AUTH" "$BASE_URL/api/gateway/runs/<run_id>/artifacts"

List artifacts visible to a session:

curl -sS -H "$AUTH" "$BASE_URL/api/gateway/sessions/sess-1/artifacts"

Browse server workspace files/folders:

curl -sS -H "$AUTH" \
  "$BASE_URL/api/gateway/files/list?path=&include_directories=true&limit=200"

Optional filters: - path: browse a specific workspace folder or mount alias. - recursive=true - family=image|video|audio|document|text|code|json|archive|other - extensions=png,jpg or newline-separated values - query=substring - max_depth=<n>

Search artifacts across Gateway storage:

curl -sS -H "$AUTH" \
  "$BASE_URL/api/gateway/artifacts/search?scope=all&artifact_kind=image&query=logo&tags=pin_id=image&include_stats=true&limit=500"

scope can be all, session, or run. Use session_id with scope=session and run_id with scope=run; omit both for scope=all. Search responses carry the row fields and also include artifact_envelope_v1, a normalized projection of Runtime-owned descriptors, access stats, and Gateway action links.

Useful query parameters: - artifact_kind: UI-oriented kind filter. Comma-separated values match semantic_kind, render_kind, or modality; generic audio means unclassified audio and does not match canonical voice, music, or sound. Single canonical kinds such as music, voice, image, markdown, or json map to Runtime catalog filters. Multi-kind unions are supported, but may be Gateway post-filters until Runtime exposes OR filters. - semantic_kind / render_kind: canonical descriptor filters when the caller wants the two dimensions separately. - modality, content_type, workflow_id, node_id, created_after, created_before, and tags: server filters for indexed descriptor fields. - query: case-insensitive metadata search. Gateway may post-filter this field when Runtime cannot index it directly. - include_stats=true: include exact stats.total, byte totals, and facet counts for the selected server-side filter set, independent of limit. - limit, offset, and cursor: bounded paging. The default Runtime Explorer page size is 500; limit<=0 is bounded unless debug_unlimited=true is used by an admin/debug caller.

artifact_envelope_v1 contains normalized fields such as semantic_kind, render_kind, workflow_id, node_id, turn_id, ledger_cursor, generation, producer, media, source_refs, access, and links. Sparse producer metadata is represented as missing fields; Gateway does not invent provider/model provenance from filenames.

Generated-media artifacts created by child runs and projected into the parent run preserve Runtime descriptors and structured metadata. Direct transcription routes store transcript artifacts with source-audio refs, language/prompt hints, provider/model when available, and bounded route parameters. Descriptor-provided action links are sanitized to relative Gateway/UI links before they appear in envelopes; raw external provider URLs should be represented as trace availability or Gateway-owned trace records.

Content reads can label the access type for Runtime access stats:

curl -sS -H "$AUTH" \
  "$BASE_URL/api/gateway/runs/<run_id>/artifacts/<artifact_id>/content?access_action=preview"

Supported access actions are content, preview, and download. The shorter access=preview alias is also accepted.

Import a server workspace path into a session artifact:

curl -sS -H "$AUTH" -H "Content-Type: application/json" \
  -d '{"session_id":"sess-1","source":{"kind":"workspace_path","path":"inputs/photo.png"},"pin_id":"image"}' \
  "$BASE_URL/api/gateway/artifacts/import"

Export an artifact back into the server workspace:

curl -sS -H "$AUTH" -H "Content-Type: application/json" \
  -d '{"path":"outputs/photo.png","create_parent_dirs":true,"overwrite":false}' \
  "$BASE_URL/api/gateway/runs/<run_id>/artifacts/<artifact_id>/export"

Import and export use the same Gateway workspace policy as file helpers: workspace roots, mounted roots, ignored paths, and size limits are enforced on the server. Browser-local files should be uploaded through POST /api/gateway/attachments/upload; browser-local file paths are not interpreted as Gateway workspace paths. An upload belongs to the conversation it was uploaded in: a run of another session that references it (any shape — context.attachments, context.media, attachments, media, a bare id or one naming its owner run_id) is refused at the start door with 400 artifact_not_in_session (and by the host for in-process callers), so one conversation's attachment can never reach another's model, whatever a client sends. Artifacts a run produced keep the hand-off below: a later run, in any session of the same user, may name them by run_id. An artifact tagged shared: user is visible to every session of its owner. In hosted user-auth mode, server workspace import/export and /files/* helpers require an admin principal. They never list, read or write the gateway data folder or the account's credential folders (.ssh, .aws, Library/Keychains, ...), even when the workspace root contains them (403). Ordinary users can still upload browser-local files and list/search artifacts in their own routed runtime.

Canonical Gateway server paths use rel/path for the main workspace root and mount_alias/rel/path for approved mounts. When two allowed mounts share the same basename, Gateway emits deterministic digest-suffixed aliases so the same public path string can round-trip through /files/*, artifact import/export, and Runtime file nodes.

Durable commands (POST /api/gateway/commands)

Commands are appended to a durable inbox and applied asynchronously by the runner.

Request fields (see SubmitCommandRequest in src/abstractgateway/routes/gateway.py): - command_id: client-supplied idempotency key (UUID recommended) - run_id: target run id (or session id for some event use-cases) - type: pause|resume|cancel|conclude|emit_event|update_schedule|compact_memory|inject_guidance, or an automation.* type with run_id = the automation id (automations.md). Run commands aimed at an automation id answer 409 invalid_state. - payload: command-specific object

Pause / cancel

curl -sS -H "$AUTH" -H "Content-Type: application/json" \
  -d '{"command_id":"'"$(python -c 'import uuid; print(uuid.uuid4())')"'", "run_id":"<run_id>", "type":"pause", "payload":{"reason":"operator_pause"}}' \
  "$BASE_URL/api/gateway/commands"

Resume a paused run

curl -sS -H "$AUTH" -H "Content-Type: application/json" \
  -d '{"command_id":"'"$(python -c 'import uuid; print(uuid.uuid4())')"'", "run_id":"<run_id>", "type":"resume", "payload":{}}' \
  "$BASE_URL/api/gateway/commands"

Resume a WAITING run with a payload (WAIT resume)

When payload.payload is present, the runner interprets this as “resume a WAITING run with a durable payload”:

curl -sS -H "$AUTH" -H "Content-Type: application/json" \
  -d '{"command_id":"'"$(python -c 'import uuid; print(uuid.uuid4())')"'", "run_id":"<run_id>", "type":"resume", "payload":{"wait_key":"<optional_wait_key>", "payload":{"approved":true}}}' \
  "$BASE_URL/api/gateway/commands"

Evidence: src/abstractgateway/runner.py (_apply_command, _apply_run_control).

Emit an external event

Minimal form:

curl -sS -H "$AUTH" -H "Content-Type: application/json" \
  -d '{"command_id":"'"$(python -c 'import uuid; print(uuid.uuid4())')"'", "run_id":"<session_id>", "type":"emit_event", "payload":{"name":"chat.message","payload":{"text":"hi"}}}' \
  "$BASE_URL/api/gateway/commands"

Evidence: src/abstractgateway/runner.py (_apply_emit_event).

Automations

An automation runs a workflow on a trigger (a fixed UTC interval, or only when asked) as a durable controller run whose occurrences are ordinary child runs. The full contract, with request and response shapes, the error envelope, typed waits, attention and operations, is in automations.md.

Route Answer
GET /api/gateway/trigger-sources {items: [{id, version, label, capabilities, config_schema, event_schema, available, unavailable_reason?}]}
POST /api/gateway/automations {request_id, title?, target, trigger?, context?, policy?} → {automation_id, revision, summary}
GET /api/gateway/automations?status=&archived_only=&cursor=&limit= {items: [AutomationSummary], next_cursor, archived_automations} (older scheduled runs on the last page, legacy: true). Archived automations are left out unless archived_only=true (only them) or status names archived; archived_automations counts yours
GET /api/gateway/automations/{automation_id} {definition, active_revision, summary}
PATCH /api/gateway/automations/{automation_id} {command_id, expected_revision?, changes} → {command_id, accepted, duplicate, seq}
POST /api/gateway/automations/{automation_id}/commands {command_id, type: automation.pause|resume|run_now|stop_current|archive|unarchive|revise, payload?} → the same receipt. unarchive brings an archived automation back paused with its history (409 invalid_state when it is not archived); an archived automation accepts only unarchive
GET /api/gateway/automations/{automation_id}/occurrences?cursor=&limit= {items: [occurrence], next_cursor}, newest first
GET /api/gateway/automations/{automation_id}/attention?cursor=&limit= {items: [attention item], next_cursor}: your unseen items, oldest first
POST /api/gateway/automations/{automation_id}/seen {attention_cursor} → {attention_cursor} (moves forward only)
POST /api/gateway/automations/{automation_id}/discuss {request_id, occurrence_index, prompt} → {session_id, run_id, session_kind: "discussion"}

Errors on these routes, 401 and 403 included, always have the shape {"detail": {"reason_code", "message", "field"?, "command_id"?}} (automations.md). An occurrence's wait is answered with the resume command above; the answer's shape depends on the wait's kind (automations.md). GET /api/gateway/runs rows carry session_kind, automation_id, role, occurrence_index and legacy, plus workspace_root on turn rows (the folder the run works in; absent on sub-runs), accept session_kind=chat,discussion, and root_only=true returns conversation turns, one per occurrence (automations.md). With include_metrics=true each row also carries steps, llm_calls, tool_calls and tokens_total: the totals of that run and every sub-run below it, read from the ledger on either store backend (null when the gateway has no ledger store). AbstractCode's conversation card shows the sum of its turns' tool_calls.

Archived sessions

A conversation can be archived to take it out of the list. Archiving deletes nothing: the session's runs, ledgers and artifacts stay, and GET /api/gateway/runs?session_id=<id> still reads it (its rows then carry archived: true and archived_at).

Route Answer
POST /api/gateway/sessions/{session_id}/archive {session_id, archived: true, archived_at, archived_by, changed}; a repeat answers changed: false
POST /api/gateway/sessions/{session_id}/unarchive {session_id, archived: false, archived_at: null, archived_by: null, changed}
GET /api/gateway/runs?root_only=true leaves archived sessions out; every listing carries archived_sessions (how many sessions of the listed kind you have archived)
GET /api/gateway/runs?root_only=true&archived_only=true only the turns of archived sessions, each with archived: true

Session purpose (kind). A session is a conversation (the default) or a docs chat (a Docs assistant conversation). POST /api/gateway/runs/start takes kind: "docs" with a session_id (the kit's Docs assistant sends it); any other value, or docs without a session, is a 400. Turn listings (GET /runs?root_only=true or archived_only=true, without session_id / parent_run_id) take kind=conversation (the default), kind=docs or kind=all, so a docs chat never appears in a conversation list while staying in the same pool: readable by session_id, ledgered, archivable. Every row carries kind. The mark lives in the plane's session_kinds.json; a session without one is a conversation. A run of the shipped docs-qa workflow marks its session docs even without kind, and the first turn listing runs a one-time migration (docs_qa_bundle_v1) that marks every existing session whose turn root ran docs-qa (matched on the workflow's bundle id).

The session's owner or an admin may archive it. Each account works in its own runtime plane, so a session that belongs to another account answers 404 {"detail": {"reason_code": "session_not_found"}}. The mark and its history (every archive and unarchive) are kept in the plane's session_archive.json; each call writes a session.archived / session.unarchived audit event. Automations are archived with the automation.archive command and brought back with automation.unarchive.

Beyond the core

/api/gateway/* also includes optional operator/tooling endpoints (reports inbox, triage queue, backlog browsing + exec runner, process manager, file/attachment helpers, embeddings, voice, discovery, …).
See: maintenance.md.

Discovery endpoints (optional)

These exist to help thin clients adapt to the deployed gateway.

  • Capabilities (best-effort): GET /api/gateway/discovery/capabilities
  • Providers/models discovery (best-effort): GET /api/gateway/discovery/providers, GET /api/gateway/discovery/providers/{provider}/models
  • Tools (thin-client allowlist help): GET /api/gateway/discovery/tools. Round 12: each process-spawning tool row (AbstractRuntime's SANDBOXED_TOOL_NAMES: execute_command, shell_exec, local_helper_start, execute_python) carries sandboxed: true|false and sandbox (the state sentence, for example "Sandboxed to this run's workspaces"), and the answer carries command_sandbox (as on GET /workspace/policy)
  • Skills inventory: GET /api/gateway/skills — the abstractskill shelf with trust verdicts (roster rows {name, description, trust_level, blocked, requires_review, tree_hash, source, has_scripts, reasons}); degradations are labeled warnings, never a fabricated list. The shelf is, in order: the saved setting skills.shelf, then the legacy launch environment value (reported as such), then the gateway's own copy in <data dir>/skills/registry, then a framework checkout's shelf (only when that copy is missing), else none; the response names which as shelf_source (stored, env, seeded, checkout, none). See Skills shelf.
  • MCP server inventory: GET /api/gateway/mcp/servers — the declared registry at <data_dir>/config/mcp_servers.json ({"version": 1, "servers": [{"name", "url"?, "description"?, "auth_required"?, "tags"?}]}), served with declared fields only and probed: false (connect state/tool counts require a probe lane and are never faked).
  • Dynamic capability catalogs: GET /api/gateway/voice/voices, GET /api/gateway/audio/speech/models, GET /api/gateway/audio/transcriptions/models, GET /api/gateway/audio/music/providers, GET /api/gateway/audio/music/models, GET /api/gateway/vision/provider_models

The capabilities payload includes package presence (abstractruntime, abstractcore, abstractmemory, abstractvoice, abstractvision), existing gateway helpers (tools, visualflow, media), memory-store readiness, and AbstractCore capability plugin status for voice, audio, vision, and music.

The route paths and contract descriptors are the stable part of this surface. Catalog routes also include a stable Gateway-owned envelope:

  • catalog.contract = gateway_catalog_v1
  • catalog.version = 1
  • items = [...]

The lower-layer fields stay in the payload for compatibility. Thin clients should read catalog plus items; the route-specific fields (models, provider_models, profiles, voices) remain available.

Provider discovery also reports the resolved default provider/model when one is configured. The resolver follows request values, flow pins, and the execution-host input.text capability route; if no pair exists, the response includes default_error rather than a hardcoded local model.

It also includes a versioned thin-client contract:

  • capabilities.contracts.version: currently 1
  • capabilities.contracts.common: shared run start/list/summary/input/history, ledger, artifact, attachment, workspace, discovery, provider prompt-cache controls, and the host-visibility descriptors model_residency (including row_schema = "model_residency_row_v1" and the canonical modality_ui color map), host_state, and session_caches (see Host state and model residency). common.artifacts includes run listing/content, session artifact listing, artifact search with artifact_envelope_v1, exact stats/facets, artifact_kind UI filtering, workspace import, and workspace export descriptors when available. Permission-sensitive descriptors are principal-aware: ordinary users see admin-only workspace import/export and provider prompt-cache controls marked unavailable with admin_required metadata.
  • capabilities.contracts.common.readiness: compact Gateway-owned gateway_surface_readiness_v1 summary derived from the shared endpoint/media/ residency descriptors
  • capabilities.contracts.flow_editor: the AbstractFlow editor/runtime surface
  • capabilities.contracts.assistant: assistant-facing voice/audio/media/cache feature gates
  • capabilities.contracts.abstractcode: code-client run/history/workspace/cache feature gates

Contract booleans are intentionally conservative. Package installed=true is not the same thing as endpoint available=true; clients should branch on the versioned contract fields when enabling controls.

common.readiness is intentionally narrower than provider/backend health. It summarizes Gateway surface availability from existing descriptors, but it does not invent selected backend/provider/model truth or stable degraded-state reason codes.

Evidence: src/abstractgateway/routes/gateway.py (discovery_capabilities, discovery_providers).

AbstractFlow gateway-first editor contract

The browser editor can use AbstractGateway as its runtime and storage host.

Draft VisualFlow records:

  • GET /api/gateway/visualflows
  • POST /api/gateway/visualflows
  • GET /api/gateway/visualflows/{flow_id}
  • PUT /api/gateway/visualflows/{flow_id}
  • DELETE /api/gateway/visualflows/{flow_id}
  • POST /api/gateway/visualflows/{flow_id}/publish

Bundle inspection and editor run-schema helpers:

  • GET /api/gateway/bundles
  • GET /api/gateway/bundles/{bundle_id}
  • GET /api/gateway/bundles/{bundle_id}/flows/{flow_id}
  • GET /api/gateway/bundles/{bundle_id}/flows/{flow_id}/input_schema

Workflow description (the console's inline edit on the Workflows page):

Route Body → answer
PATCH /api/gateway/bundles/{bundle_id} {description} (at most 2000 characters; "" goes back to the file's own description) → {ok, bundle_id, owner, description, description_edited, updated_by, updated_at}

Only the owner may: a user for their own workflows, an admin for the gateway's. A user on a shared workflow gets 403 admin_required, another user's workflow is 404, a shipped workflow 409 workflow_shipped. The .flow file is never rewritten (Export keeps the original bytes); the text applies to every version and is kept next to the owner's archive file (config/workflow_descriptions.json). Each attempt is audited as workflow.description (lengths, never the text). GET /api/gateway/bundles items carry the effective description, description_edited and actions.can_edit_description; the envelope's interfaces ({<interface id>: {label, help, known}}) names every interface the listed workflows declare, readable by any signed-in account (the console's "Used by" column).

The input-schema endpoint returns a versioned payload with:

  • version
  • bundle_id, bundle_version, bundle_ref, flow_id, workflow_id
  • inputs: entrypoint input pins derived from the on_flow_start node
  • defaults: pin defaults from VisualFlow JSON
  • input_data_schema: a small JSON Schema object for the Run Flow modal

Example:

curl -sS -H "$AUTH" \
  "$BASE_URL/api/gateway/bundles/my-bundle/flows/ac-echo/input_schema"

Native-loop bundles (react / codeact / memact)

Some shipped bundles declare metadata.native_loop_factory instead of VisualFlow JSON (manifest.flows is empty). The gateway materializes an abstractagent loop at load time. Discovery uses the same bundle list endpoint — not /discovery/workflows.

Thin clients should:

  1. GET /api/gateway/bundles (authenticated).
  2. Filter entrypoints whose interfaces includes abstractcode.agent.v1.
  3. Read metadata.native_loop_factory (react, codeact, or memact) to distinguish native loops from VisualFlow agent bundles.
  4. Use each entrypoint's workflow_id (for example react-agent@0.1.0:react).
  5. Start runs with POST /api/gateway/runs/start and bundle_id / flow_id from the bundle listing (for example react-agent + react).

Native-loop entrypoints do not ship VisualFlow JSON. The gateway serves a versioned input-schema stub (prompt required; provider and model optional) from GET /api/gateway/bundles/{bundle_id}/flows/{flow_id}/input_schema. Headless clients may also pass those fields without fetching the schema.

The shipped react-agent@0.1.0 bundle is built by scripts/build_react_agent_bundle.py and force-included in the wheel. A running gateway process must restart (or call bundle reload) after the file lands on disk before /bundles lists it.

Run history bundle (GET /runs/{run_id}/history_bundle)

Thin clients should prefer this endpoint over stitching ledger, session, and artifact endpoints. The export is owned by AbstractRuntime; the gateway forwards query parameters and returns the bundle JSON unchanged (including in-band degradations).

Query parameters:

Parameter Default Notes
include_subruns true Descendant runs in the bundle tree
include_session false Root session turn list
session_turn_limit 200 Cap when include_session=true
ledger_mode tail tail or full
ledger_max_items 2000 Per-run ledger cap when ledger_mode=tail
detail full full (complete payloads) or replay (transcript-fold projection)

detail=replay drops request-side payloads and observability paths the transcript fold never reads; runtime marks each omission with $omitted inside ledger records. Use it for session replay and thin-client folds — it is much smaller than full (gzip helps further; send Accept-Encoding: gzip).

warnings (always present, may be empty): typed degradations the export survived instead of failing silently. Each entry is an object with at least code and detail; many include run_id. Known codes today:

Code Meaning
subtree_discovery_failed Could not list child runs; bundle covers root only
subtree_truncated Descendant discovery hit the run cap
ledger_read_failed Ledger for a run id could not be read
torn_rows_skipped Corrupt/unparseable ledger lines skipped
ledger_tail_window Ledger truncated to ledger_max_items (tail mode)
input_data_offload_failed Input-data artifact reference could not be resolved

Clients must surface non-empty warnings to the operator — a bundle that "looks complete" but carries warnings may be missing subruns, ledger tail, or offloaded input data.

Session history bloc (GET /sessions/{session_id}/history/bloc)

Returns one cursor-bounded bloc of root session turns, each with an inline history_bundle export — one round-trip instead of N per-turn bundle fetches. Resume pagination uses an ISO created_at cursor in the before query parameter (never turn-count offsets). The turns are the ones GET /api/gateway/runs?root_only=true lists for the session: parent-less runs plus automation occurrences (a retried occurrence once, as its last attempt), never automation controllers (automations.md).

Query parameters:

Parameter Default Notes
before (omit) ISO-8601 cursor; only turns strictly before this timestamp
limit 5 Max turns in this bloc (1–50)
detail replay Forwarded to each turn's bundle export (full | replay)
include_subruns true Per-turn bundle tree
ledger_mode tail tail or full
ledger_max_items 2000 Per-turn ledger cap when ledger_mode=tail
include_drafts false Include draft-test root runs

Response fields: session_id, cursor_before (echo of before), cursor_after (oldest turn returned — pass as the next before), older_remaining, warnings, and turns[] (run_id, created_at, status, bundle or error).

The editor observes runs with the core lifecycle endpoints above: /runs/start, /runs/{run_id}, /runs/{run_id}/ledger, /runs/{run_id}/ledger/stream, /runs/ledger/batch, /runs/{run_id}/input_data, /runs/{run_id}/history_bundle, and /runs/{run_id}/artifacts.

Optional multimodal scope

Current direct Gateway endpoints: - POST /api/gateway/runs/{run_id}/voice/tts - POST /api/gateway/runs/{run_id}/voice/tts/stream - POST /api/gateway/runs/{run_id}/audio/transcribe - POST /api/gateway/runs/{run_id}/images/generate - POST /api/gateway/runs/{run_id}/images/edit - POST /api/gateway/runs/{run_id}/images/upscale - POST /api/gateway/runs/{run_id}/videos/generate - POST /api/gateway/runs/{run_id}/videos/from_image - POST /api/gateway/runs/{run_id}/music/generate - GET /api/gateway/voice/defaults - GET /api/gateway/voice/voices - GET /api/gateway/audio/speech/models - GET /api/gateway/audio/transcriptions/models - GET /api/gateway/audio/music/providers - GET /api/gateway/audio/music/models - GET /api/gateway/vision/provider_models - GET /api/gateway/vision/adapters

/voice/defaults is the one answer to "which engines speak and listen by default":

{"tts": {"route": "output.voice", "configured": true, "provider": "supertonic", "model": "supertonic-3", "voice": "M3"},
 "stt": {"route": "input.voice", "configured": true, "provider": "faster-whisper", "model": "large-v3"},
 "source": "capability_defaults"}

A route the administrator has not set reads {"configured": false, "provider": null, "model": null, "note": "No gateway default is set for …"}. A synthesis or transcription request that names no provider and no model runs exactly these routes, so apps show "Gateway default · supertonic / supertonic-3" from here — never from the voice catalog's engine-side fields. The catalog repeats the answer (gateway_defaults; active_tts_provider / active_stt_provider = the configured routes, absent when unset). /audio/transcribe returns the route that ran (provider, model) and duration_ms; a language hint skips the engine's language detection (faster-whisper large-v3 on an Apple-Silicon CPU: ~25 s → ~9 s for a 4 s clip).

/voice/tts returns a durable audio artifact after synthesis. /voice/tts/stream returns JSON Lines stream events for progressive playback when discovery advertises capabilities.contracts.assistant.voice.tts.streaming=true; successful streams still finish with a Runtime-owned child-run audio artifact. The stream is real streaming: the engine splits the text at sentence boundaries (first segment one short clause), each segment is sent as soon as it is synthesised while the next one is synthesised, and the setup runs off the event loop. The terminal done event's metrics carry ttfb_s, rtf, device and, on the CPU, device_reason.

Stream lines (JSON Lines): runtime_start (run_id = the run named in the path, child_run_id = the id the outcome will be recorded under, schema: abstractruntime.tts_stream.start.v2), engine events (start, audio with sequence and audio_b64), then one terminal line: done (with audio_artifact and child_run_status), cancelled or error (error = the sentence; watchdog_timeout: true when the engine went silent past ABSTRACTGATEWAY_VOICE_TTS_TIMEOUT_S). When the engine is busy (another stream, or the model using the machine) and its setup takes more than a second, the first line is {"type": "queued", "message": "Waiting for the voice engine: …"}; speech follows when the engine is free. A request past the voice synthesis bound (ABSTRACTGATEWAY_VOICE_MAX_CONCURRENCY) waits one second, then gets 503 "Read aloud is busy: N voice syntheses are already running on this gateway …".

Reading aloud never changes the run's state: no wait is created while audio streams, and the child run is recorded already completed (done, or with errors[0].code cancelled/stream_error) when the stream ends. A stream cut by a gateway restart leaves no run behind; a wait left by a gateway up to 0.13.0 is closed at startup with errors[0] = {"code": "interrupted", "message": "Read aloud was interrupted because the gateway restarted; …"}.

The catalog endpoints proxy AbstractCore Server routes when ABSTRACTCORE_SERVER_BASE_URL is configured. Gateway uses explicit Core auth settings for that hop and never reuses the Gateway bearer token as a Core/provider secret. Without a configured Core server, the voice/model routes return bounded static descriptors from Gateway and capability-package environment variables.

Each route adds:

  • catalog: Gateway-owned route metadata (contract, version, kind, scope, route_source, optional upstream_source, and route filters)
  • items: one canonical primary array for thin clients

Examples:

  • Voice provider listings (/voice/voices?providers_only=true, /audio/speech/models?providers_only=true, /audio/transcriptions/models?providers_only=true) always list the cloud providers openai and openai-compatible; their items carry needs_key, key_source (environment | providers | null), state and reason, and cloud_providers repeats them. A key counts when it is in the environment or saved through the Providers screen. Local engines are listed when AbstractVoice reports their runtime installed; unavailable_providers (per kind: {provider, code: runtime_missing|model_not_downloaded|not_configured, reason}) and unavailable_reason say why others are missing, and a listing filtered to a cloud provider without a key says where the key goes.
  • /voice/voices: items contain voice/profile records with id, label, optional provider, optional model, and voice_kind
  • /audio/*/models: items contain model records with id, label, optional provider, optional tasks, and optional parameters
  • /audio/music/providers and /discovery/providers: items contain provider records with id, label, and provider

Generated images are available through Runtime workflows when a compatible image backend is installed and configured. Gateway also exposes a direct image generation endpoint that uses the Runtime/Core output-selector contract rather than a provider-specific image client. The route creates a durable child run, stores the generated image as a run artifact, and returns event_name="abstract.progress" so thin clients can stream the child-run ledger for progress:

  • run_id, request_id, prompt
  • optional provider, model, size, width, height, format, batch count / n, seeds, and ordered lora_adapters
  • image_artifact: first generated image for compatibility
  • image_artifacts: full ordered image artifact list for batch generation

size, width, and height are optional passthrough request overrides. Do not inject a client-side default size. Different image providers/models accept different size sets; when the client leaves dimensions unset, Runtime/Core lets the configured backend use its default or auto behavior.

If the active workflow runtime already has an AbstractCore LLM client, the route uses it. For tools-only workflows, the route can create a direct Runtime/Core client from request provider/model or the execution-host capability route default. Unsupported or unconfigured deployments return a structured ok=false response instead of a failed run.

Gateway also exposes a direct image-edit sibling route:

  • POST /api/gateway/runs/{run_id}/images/edit

The request uses a source image_artifact, optional mask_artifact, the same provider/model and image backend selectors as image generation, plus optional batch count / n, seeds, and ordered lora_adapters, and returns an artifact-backed edited image. Batch responses also return image_artifacts. Thin clients should feature-detect it from capabilities.contracts.flow_editor.media.edited_image or capabilities.contracts.assistant.media.edited_image. It uses the same child-run abstract.progress progress contract as direct image generation.

Gateway also exposes a direct image-upscale sibling route:

  • POST /api/gateway/runs/{run_id}/images/upscale

The request uses a run-visible source image_artifact, optional provider/model selectors, and optional upscaler controls such as scale, resolution, softness, seed, quantize, and vae_tiling; resolution may be a shortest-edge integer or a scale factor such as 2x. Thin clients should feature-detect it from capabilities.contracts.flow_editor.media.upscaled_image or capabilities.contracts.assistant.media.upscaled_image, list models with GET /api/gateway/vision/provider_models?task=image_upscale, and stream the returned child-run ledger for abstract.progress events.

Generated music follows the same direct child-run pattern. Thin clients should discover it from capabilities.contracts.flow_editor.media.generated_music or capabilities.contracts.assistant.media.generated_music, list providers/models from the music catalog routes, and treat the returned child_run_id plus music_artifact as the durable output handle.

Generated video also follows the direct child-run pattern:

  • POST /api/gateway/runs/{run_id}/videos/generate uses the Runtime/Core output.modality=video / task=text_to_video contract and accepts optional batch count / n, seeds, ordered lora_adapters, and flow_shift.
  • POST /api/gateway/runs/{run_id}/videos/from_image accepts a run-visible source image_artifact, accepts the same optional batch/adapter/video control fields, and uses task=image_to_video.
  • Thin clients should discover these routes from capabilities.contracts.flow_editor.media.generated_video and capabilities.contracts.flow_editor.media.image_to_video (or the matching assistant.media.* entries), use GET /api/gateway/vision/provider_models?task=text_to_video|image_to_video for model catalogs, use GET /api/gateway/vision/adapters for compatible installed adapter catalogs, stream the returned child_run_id ledger for abstract.progress events, and read video_artifacts when batch generation is requested.

STT and listen contract notes:

  • POST /api/gateway/runs/{run_id}/audio/transcribe accepts a run-visible audio_artifact plus optional language, prompt, response_format, temperature, format, provider, and model hints.
  • capabilities.contracts.flow_editor.voice.stt and capabilities.contracts.assistant.voice.stt point to that upload route.
  • capabilities.contracts.flow_editor.voice.listen and capabilities.contracts.assistant.voice.listen are host-capture contracts, not a live microphone socket. They tell higher apps to capture locally and emit an event or upload the resulting audio artifact.

KG memory

POST /api/gateway/kg/query queries the configured AbstractMemory TripleStore. Gateway resolves the store through:

  • ABSTRACTGATEWAY_MEMORY_STORE_BACKEND=lancedb|memory (sqlite when the installed AbstractMemory build exposes SQLiteTripleStore)
  • ABSTRACTGATEWAY_MEMORY_STORE_PATH
  • ABSTRACTGATEWAY_MEMORY_REQUIRE_VECTOR

Structured queries work with LanceDB and in-memory stores. SQLite also works when the installed AbstractMemory build exposes SQLiteTripleStore. Semantic query_text requires a vector-capable backend plus the execution-host embedding.text route; SQLite returns a clear 400 instead of pretending to support semantic recall.

Capability discovery reports KG memory as available when AbstractMemory is installed and the configured backend can be resolved. A fresh persistent store does not need to exist yet; empty-store structured queries return an empty result rather than making Flow authoring nodes unavailable.

Models and engines

The gateway serves AbstractCore's models and engines payloads unchanged, under /api/gateway. The bodies and payloads are the same as AbstractCore's own /acore/* routes; abstractcore and abstractgateway render them with the same screens.

Method and path Access Body / query Returns
GET /host/profile user refresh=1 host_profile_v1
GET /engines user probe=1 gateway_engines_v2 rows (AbstractCore's detection plus the install plan and actions) with install_allowed and install_policy; see engines.md
GET /engines/{id} user probe=1 one engine row plus install_allowed; 404 for an unknown id
POST /engines/{id}/install admin {"dry_run": bool, "force": bool, "location": "auto"\|"user"\|"system"} an engine_install_job_v1 job (user-level first; pauses in needs_admin / needs_tools), see engines.md
GET /engines/jobs, GET /engines/jobs/{id} user engine install jobs
POST /engines/jobs/{id}/continue, /cancel admin {"action"?} the job
POST /engines/{id}/start, /stop admin Ollama / LM Studio server state
GET /models/catalog user q, engine, fits=1, hub=1, tag (repeatable) model_catalog_v1
GET /models/installed user provider models_installed_v1
POST /models/download admin {"provider", "artifact", "dry_run", "expected_bytes"?} or {"recommended": true} {"ok": true, "job": {...}}; with recommended, {"ok": true, "recommended": true, "jobs": [...], "group": {...}}
GET /models/download/{job} user {"ok": true, "job": {...}} (a grp_... id returns the parent job)
GET /models/downloads user {"ok": true, "jobs": [...]}, newest first, parents included
POST /models/download/{job}/cancel admin none {"ok": true, "job": {...}}; stops the transfer within about a second; a grp_... id cancels every running child; 404 when unknown
GET /models/downloads/stream user job_id, until_idle=1 Server-Sent Events of the same dicts, see model-downloads.md
POST /models/delete admin {"provider", "artifact", "dry_run": bool, "force": bool} host_job_v1 (kind delete)
POST /models/delete-download admin {"provider", "artifact", "dry_run": bool} (the catalog's own names) model_download_delete_v1, see "Delete a download" below
GET /jobs user kind, status {"schema": "host_jobs_v1", "jobs": [...], "generated_at"}, newest first
GET /jobs/{id} user host_job_v1; 404 when unknown
POST /jobs/{id}/cancel admin none (an empty {} is accepted) host_job_v1; 404 when unknown

Empty query values (q=, engine=) mean "no filter". probe, fits and hub accept 1/0 and true/false.

Catalog artifacts. model_catalog_v1 is AbstractCore's payload, served unchanged (field reference: AbstractCore docs/models.md, "The catalog"). Besides quant (the artifact's own label, lowercased, or null) and bits (effective bits per weight), every artifact carries quant_class, one of 2bit, 3bit, 4bit, 5bit, 6bit, 8bit, 16bit, full, unknown, for filtering by quantization: q4_k_m, 4bit, mxfp4 and oq4e are 4bit; q8_0 and 8bit are 8bit; bf16 and f16 are 16bit; f32 is full. quant_class_source is stated (the reference names its quant), assumed (a bare Ollama tag such as qwen3.5:9b or LM Studio id: the class of the engine's default build, which the fit estimate assumes too) or null (no quant information; the class is unknown). options holds the route options a recommendation copies with the artifact ({} for most); companions lists repos downloaded with it (an MLX build's MTP drafter, from AbstractCore's drafter registry; [] for most), companion_bytes is their size, and download_bytes already includes it; note is one sentence about the build. On Apple silicon the text rows pre-select the memory tier's MLX build and exactly one text row is the starter.

Jobs. A host_job_v1 has schema, job_id, kind (download | delete | engine_install), status (queued | running | completed | failed | cancelled), provider, artifact, engine, percent, downloaded_bytes, total_bytes, message, log_tail, command (the exact argv), dry_run, started_at, finished_at, error (a string or null), joined, result and cli_equivalent, which names the abstractgateway command that does the same thing. A dry run finishes before the POST returns. On the /models/download routes the job also carries job (the id), events, host_status, reports queued as running, and counts joined including the first request.

Download progress. A download job also carries state (queued | resolving | downloading | verifying | installing | done | failed | cancelled | stalled), bytes_done, bytes_total, size_unknown, size_note, bytes_per_second, eta_s, updated_at, files ([{name, bytes_done, bytes_total, state}]), current_file, a one-sentence message, the tool's own detail, and transitions. "Use recommended defaults" ({"recommended": true}) returns one parent job (kind: "download_group", id grp_...) whose bytes, percent, speed and time left add up its children. The full contract, one real example per state and what each source reports: model-downloads.md.

Refusals share one body: {"ok": false, "status", "reason"?, "message", "detail", "error": {"message", "type"}, ...}.

Status When
400 invalid provider or artifact missing
403 refused / not_allowed a real engine install while allow_engine_install is off (configuration.md); install_policy says why
403 the caller is not an admin (every POST above)
404 not_found unknown job id, engine id, or a model that is not installed
409 busy an engine install is already running (job is the running one)
409 refused the engine is not supported here or has no install command (install is the plan), or a delete is blocked (delete_blockers: loaded, shared_cache:…, unknown_location, engine_not_running, remote_engine; force: true overrides the first two)
501 unsupported / abstractcore_too_old the installed AbstractCore is too old for these routes; required, installed and missing name what to upgrade
503 unavailable AbstractCore is not installed

Delete a download (POST /models/delete-download, the console's Models-page Delete). Synchronous; removes one downloaded artifact with its engine's own mechanism: Ollama's DELETE /api/delete (what ollama rm does), the Hugging Face / MLX cache folder of the repo, or, for org/repo:QUANT (a llama.cpp GGUF quant), only that quant's files (the repo goes when it was the last model file set). dry_run: true deletes nothing and answers the exact bytes for the confirmation. Answers:

Status Body
200 {"schema": "model_download_delete_v1", "ok": true, "status": "planned" \| "deleted", "provider", "artifact", "freed_bytes", "paths", "command", "also_used_by", "presence": "installed" \| "absent", "message"}
409 refused reason: resident (this gateway or the engine has it loaded: "Unload it first"), locked ("Unlock and unload it first"), downloading (its download is still running), managed_elsewhere (LM Studio: its CLI has no remove command, delete it in LM Studio), engine_not_running, unknown_location, remote_engine; message and fix are sentences to show as they are
404 not_found reason: "not_downloaded"
502 failed reason: "engine_failed": the engine did not delete; command names what ran

A build whose files AbstractCore files under the sibling engine of the shared Hugging Face cache (MLX vs Hugging Face) is deleted all the same, and also_used_by names that engine. Every real delete writes model.download_deleted (provider, artifact, actor, freed_bytes, paths) to <data_dir>/audit_log.jsonl; every refusal (a dry run's too) writes model.download_delete_refused with its reason and dry_run.

Example:

curl -s -H "Authorization: Bearer $TOKEN" "$GW/api/gateway/models/catalog?q=qwen3&fits=1" | jq '.rows[0].artifacts[0].fit'
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
  -d '{"dry_run": true}' "$GW/api/gateway/engines/ollama/install" | jq '.command, .cli_equivalent'

Host state and model residency

Gateway exposes a host-level view of the execution machine — memory, GPU, resident models, and session prompt caches — so consoles and agents can render an "agentic OS" panel from one API surface.

Read endpoints (any authenticated principal):

  • GET /api/gateway/host/state — one-call host snapshot
  • GET /api/gateway/host/metrics/memory — host memory snapshot
  • GET /api/gateway/host/metrics/gpu — GPU utilization probe
  • GET /api/gateway/models/loaded — model residency listing
  • GET /api/gateway/models/context_estimate — context/KV memory estimate for a provider+model
  • GET /api/gateway/sessions/prompt_cache — session prompt-cache enumeration

Mutation endpoints (admin principal required):

  • POST /api/gateway/models/load — load (and by default pin) a model runtime
  • POST /api/gateway/models/unload — unload a model runtime
  • POST /api/gateway/models/lock — lock a resident model against unload
  • POST /api/gateway/models/unlock — release a model-residency lock
  • POST /api/gateway/models/download — fetch model weights onto the host
  • POST /api/gateway/sessions/{session_id}/prompt_cache/clear_all — clear every runtime-minted prompt cache for a session

Reads are visibility every authenticated client needs; mutations spend shared host resources and stay operator acts. Anonymous requests are rejected on all of these routes, like every other /api/gateway/* path.

GET /host/state

One snapshot with memory, gpu, models, and session_caches sections:

{
  "ok": true,
  "ts": 1787857000.0,
  "memory": {"ram": {"...": "..."}, "process": {"rss_bytes": 140443648}, "device": {"backend": "metal", "allocated_bytes": 0, "...": "..."}},
  "gpu": {"supported": true, "source": "ioreg", "gpus": [{"name": "...", "utilization_gpu_pct": 0.0}]},
  "models": [{"runtime_id": "...", "provider": "...", "model": "...", "resident": true, "...": "..."}],
  "session_caches": [],
  "totals": {"models": 2, "models_resident": 1, "model_bytes": 3109915433, "session_caches": 0, "session_cache_bytes": null},
  "degraded": [],
  "row_schema": "model_residency_row_v1"
}
  • models rows use the frozen model_residency_row_v1 schema described below; session_caches relays the runtime facade's cache rows verbatim.
  • Every section is independently best-effort and the route never returns a
  • A missing facade method or a failed probe nulls that section and names it in degraded; a reasons map (present only when non-empty) says why.
  • The gpu section keeps its in-band {"supported": false, "reason": "..."} payload when the probe answers but reports no support; it still counts as degraded.
  • totals.model_bytes sums the known size_bytes values and is null when no row reports a size; totals.session_cache_bytes behaves the same over the cache rows' bytes.
  • memory.device relays AbstractCore's accelerator figures, including process_held_bytes (what this gateway process holds across its model libraries: MLX, llama.cpp, transformers, embeddings) and process_held_basis, how that figure was measured (metal_device_counter, cuda_device_counter…, or sum: of the libraries' own counts). memory.resident / memory.held name the models that hold it (backend, model, number of holders) when the runtime reports them.
  • residency_diagnostics is always present ({} when the runtime reports nothing): the dict GET /models/loaded returns as diagnostics, with pending_ejects (models that will be ejected when their in-flight call ends) and last_switch_ejects (what the last default-model switch ejected, kept or failed to eject, with the reason). The consoles and the tray show these as "Will eject X when the in-flight call ends", "X: eject failed: reason".
  • totals.models counts every known row — configured / cached rows included — while totals.models_resident (additive) counts only rows with resident: true. Clients that display "N loaded" must read models_resident: default ≠ loaded, and presenting configured capability defaults as loaded is exactly the lie this field removes.
  • When the runtime memory snapshot reports a host identity, the response also carries a top-level host object (the identity facts of the machine the snapshot describes). The block is omitted when the runtime does not report one. Together with the per-row host_id/host_name fields below, this is the seam a multi-machine resource pool would aggregate on; one gateway binds one runtime host, and the pool design is proposed in backlog 0093.

GET /host/metrics/memory

Returns {"ok": true, "supported": true, ...} plus the snapshot sections: ram (total/available/used bytes and percent), process (rss_bytes), and device (backend, allocated_bytes, total_bytes, free_bytes). When the runtime host facade does not expose a memory snapshot, the route answers 200 with {"ok": true, "supported": false, "reason": "..."} — the same degraded style as GET /host/metrics/gpu.

How to compare memory measurements: process.rss_bytes and device.allocated_bytes are different axes. In-process device backends (for example MLX on Metal) return freed buffers to the process heap and the operating system may retain those pages, so process RSS does not shrink when a model unloads. Use device.allocated_bytes to verify that an unload freed device memory; use ram and process for overall host pressure.

Model residency (/models/loaded, /models/load, /models/unload)

GET /models/loaded lists the model runtimes the host knows about, with optional task, provider, model, and base_url query filters. The response keeps the raw runtime records in models and adds a normalized rows array in the frozen model_residency_row_v1 schema (named by row_schema), so thin clients do not need per-provider alias tables.

Each model_residency_row_v1 row has exactly these fields (unknown values are null, never guessed):

runtime_id, task, provider, model, source, resident, state, pinned, default, size_bytes, size_vram_bytes, expires_at, context_length, loaded_at, last_used_at, locked, lockable, modalities, calibrated_context_length, context_calibrated, host_id, host_name, details

The schema is additive-tolerant and keeps the model_residency_row_v1 name as optional fields are added; treat fields beyond the original 16 as optional. The lock/calibration/host fields mean:

  • locked / lockable — tri-state booleans: whether the model is locked against unload, and whether this runtime supports locking it at all.
  • modalities — list of modality strings when the runtime reports one (null otherwise, including when the value is not a clean string list).
  • calibrated_context_length / context_calibrated — the measured usable context length and whether it came from calibration rather than metadata.
  • host_id / host_name — identity of the machine serving the model, stamped by the runtime that reported the row (see the host block note under GET /host/state above).

Residency truth is provider-first: provider_resident / provider_loaded booleans in the source record outrank the runtime-lease booleans resident / loaded, because a runtime can hold a lease on a model the provider has already evicted. A loaded-looking state string (provider_loaded, loaded, resident) can confirm residency, but a state string is never proof of absence — with no boolean present and no loaded-like state, resident stays null. details preserves the raw record for fields outside the schema.

POST /models/load accepts task (default text_generation), provider, model, optional provider options, pin (default true), base_url, timeout_s, and lock (default false) — with lock: true a successful load is immediately locked against unload, and the lock outcome is reported additively under lock in the response (a lock failure or a runtime without lock support never turns the successful load into a failure). POST /models/unload selects the runtime by runtime_id or by task/provider/model. Both relay Runtime's host facade and return the normalized residency response (operation, affected records, and in-band ok=false errors instead of opaque failures). A load response carries loaded_new (true when the call loaded the model, false when it was already in memory) and the matching action (loaded / already_loaded).

A local image or video model loaded with task: "image_generation" (or "video_generation") serves the following requests for that model on the loaded pipeline, in the gateway process; a model that was not loaded runs each request in an isolated worker process that loads it for that request.

One unload failure gets a real status code: when the target model is locked, POST /models/unload answers HTTP 409 with the normalized refusal payload as the body (ok: false, error: "model_locked", plus whatever detail the runtime included), so clients can offer force-unload or point at /models/unlock. Sending "force": true in the unload request unloads the model despite the lock. Every other unload outcome stays in-band at 200.

Model locks and context estimates

  • POST /models/lock and POST /models/unlock (admin) pin a resident model against unload and release that pin. The body selects the target like unload does: runtime_id, or provider + model, with optional base_url and timeout_s. Lock requires provider-verified residency: a configured or merely-warm model refuses with an error: "model_not_resident" payload (load it with lock: true instead); unlock always works, even for a since-evicted model, so locks are never stranded. Rows report lockable so clients know whether a lock can work, and locked so they can render the current state.
  • GET /models/context_estimate?provider=&model=&context_length= (any authenticated principal) relays the Runtime host facade's context/KV memory estimate for a provider+model. provider and model are required; context_length is optional and must be >= 1 (schema-rejected with 422 otherwise). The estimate reports its confidence in-band — calibrated, estimated, or unknown — alongside facade fields such as predicted_max_context (the context that fits beside the weights), the tri-state fits_weights / fits_requested_context split, budget_bytes, est_kv_bytes, and notes (which state the budget basis and reserve). The estimate is advisory only — no load path gates on it.

Like the other host-facade relays, these routes never 500 on capability gaps: a runtime without the method answers 200 with ok: false, available: false, and code = "model_residency_unavailable" (lock/unlock) or code = "context_estimate_unavailable" (estimate); facade exceptions use the matching *_error codes.

Session prompt-cache enumeration

  • GET /api/gateway/sessions/prompt_cache?session_id=<optional>
  • POST /api/gateway/sessions/{session_id}/prompt_cache/clear_all (admin)

The list route enumerates the prompt caches the runtime actually minted. Each cache row carries the provider/model/runtime identity, byte and token counts, and stamped attribution metadata (session_id, run_id, workflow_id, node_id). Omit session_id to list every session's caches. This enumeration lane is the recommended way to observe and reclaim session cache state: unlike the identity-derived session lifecycle endpoints described under the prompt-cache control plane below, it cannot miss caches whose keys the gateway never derived.

clear_all unloads every runtime-minted cache for one session in a single call. It requires an admin principal because it accepts any session id and clears real provider cache state; the identity-derived, caller-scoped session lifecycle endpoints remain user-level.

When the runtime facade does not expose enumeration, both routes answer 200 with ok=false, available=false, and code="session_caches_unavailable" (facade errors use code="session_caches_error"); the list route always carries a caches array and clear_all always carries cleared and count.

Discovery descriptors

GET /discovery/capabilities advertises this surface under capabilities.contracts.common:

  • model_residency: endpoints (loaded, load, unload, lock, unlock, context_estimate), the per-task support map, row_schema = "model_residency_row_v1", and modality_ui — the canonical modality color map ({version: 1, colors: {...}}, one {color, label} entry per residency task plus an unknown fallback) so every client renders the same modality palette instead of hardcoding its own. It is a rendering contract, not a runtime capability, so it is served even when the runtime facade is absent.
  • host_state: endpoints (state, memory, gpu) plus memory_available. The state route itself always answers; per-section truth lives in the payload's degraded list.
  • session_caches: endpoints (list, clear_all) plus available, reflecting whether the runtime facade supports cache enumeration.

Evidence: src/abstractgateway/routes/gateway.py (host_state, host_memory_metrics, model_residency_loaded, model_residency_lock, model_context_estimate, session_prompt_caches_list) and src/abstractgateway/security/authorization.py (route-family policy).

About (GET /api/gateway/about)

Public (no sign-in): which versions this gateway runs, for About screens.

{"abstractframework": "0.5.0", "abstractgateway": "0.7.1",
 "packages": {"abstractcore": "2.18.0", "abstractruntime": "0.7.0", "abstractskill": "0.3.0"}}

abstractframework is the AbstractFramework release the installer recorded for this gateway (FRAMEWORK_VERSION in the data directory's bootstrap.env) when there is one, else the version of the abstractframework package installed beside the gateway, else null (not installed). The installer puts only abstractgateway[...] in the gateway's environment, so its record is the one that names the release. Versions only: no paths, host names or settings.

Host control (pause, desktop tray, restart, update)

The process's own controls — the surface behind the desktop tray icon and the console's Gateway card (see tray.md). Reads are available to any authenticated principal; writes require an admin principal.

  • GET /api/gateway/host/runner — execution state:
{"ok": true, "paused": true, "paused_at": "2026-09-05T06:38:26+00:00", "paused_by": "default/admin", "reason": "meeting",
 "inflight_ticks": 0, "scope": "workflow runner", "runner_in_process": true, "step_gate_supported": true,
 "runners": [{"status": "paused", "...": "..."}], "degraded": false,
 "capabilities": {"restart": true, "shutdown": true, "reason": null, "update_job_running": false}}
  • POST /api/gateway/host/pause (body {"reason": "..."} optional) and POST /api/gateway/host/resume — both answer the payload above.
  • GET /api/gateway/host/metrics/live — {gpu, memory, runner} in one call (1 s caches); gpu/memory carry the same in-band supported shape as /host/metrics/gpu and /host/metrics/memory.
  • GET /api/gateway/host/runs?limit=25&window_hours=24 — the runs on this host across every data plane (admin; /runs answers only for the calling principal's plane). {ok, items: [{run_id, workflow_id, label, status, activity, role, created_at, updated_at, ledger_len, plane, started_epoch, updated_epoch, observer_path}], count, active_count, has_more, planes, skipped_entity_planes?, warnings?}. Items are TURN ROOTS (the runtime's is_turn_root: parent-less runs that are not automation controllers, plus automation occurrences), in this order: every active root first — activity: "running" (the root or any run below it is running), then "waiting" (waiting for a person or an event, updated inside the window) — then roots that finished inside window_hours of their last update (activity: "done"), newest first. Active rows are never windowed and never cut by limit (active_count counts them); limit bounds the finished rows and has_more says more finished ones exist. label decodes a catalog workflow's internal id (__catalog__v2__…<base64>) to the name an operator uses; observer_path is the run's Observer page (/apps/observer/#run/<id>, open it signed in through POST /api/gateway/apps/observer/open {path: "/#run/<id>"}). The gateway's own bookkeeping runs (__-prefixed, but never a catalog id) are excluded.
  • GET /api/gateway/host/tray — {dependencies_installed, install_hint, decision: {start, reason, hint}, supervisor: {running, ready, pid, exit_code, failure, log_path}, can_control}. There is no setting: the icon is shown whenever this process and this desktop can hold it.
  • POST /api/gateway/host/tray/show — retry the helper now (409 when this process cannot). No hide counterpart, by design.
  • POST /api/gateway/host/restart, POST /api/gateway/host/shutdown — {"ok": true, "restart": true, "requested_by": "...", "reason": "..."}; 409 with a plain reason when unsupported (--reload, embedded server, an update is installing).
  • GET /api/gateway/host/start-at-login, PUT /api/gateway/host/start-at-login {enabled, replace_other?} — admin. Would this gateway start at the next login, and can it be changed from here: {enabled, state: on|off|broken|other, mechanism, mechanism_label, can_change, reason, summary, problems, location, platform, experimental[, other_data_dir]}. The mechanism is a LaunchAgent (macOS), a systemd user unit or, without a systemd user manager, a desktop autostart entry (Linux), or a HKCU\…\Run value (Windows). can_change: false carries the reason (for example a Linux server with neither a systemd user manager nor a desktop session). PUT registers for the next login without starting a second gateway, or unregisters without stopping this one, and answers the read-back under start_at_login; 409 when it cannot change here or when another gateway's registration would be replaced without replace_other: true; 500 when the change does not read back. The same switch as the tray's Start at login and abstractgateway service enable|disable.
  • GET /api/gateway/host/update, POST /api/gateway/host/update/check, POST /api/gateway/host/update/start (admin) — all three answer {current, install: {kind, path, upgradable, reason, command, display_command, extras, data_dir, framework_version}, check: {source, latest, update_available, offline, checked_at, error, release}, job: {state, command, log_tail, exit_code, error, message, restart_recommended, version_before, version_after, installer, framework_before, framework_after, changes, other_changes}, restart_pending, update}. install.kind is installer for an AbstractFramework installer install (check.source: framework-release, check.release: {version, gateway_version, installed, commit, manifest_url, installer: {url, commit_url, sha256, size}}); other kinds compare with PyPI (source: pypi). update is the rendering every client shows: {status, line, hint, offer, action: {label, confirm, command, source, installer_sha256} | null, checked_at}, status one of not_checked, available, not_possible, up_to_date, offline, error, running, installed, no_change, failed. start takes {installer_sha256} (the action's; required for an installer install, optional otherwise) and answers 409 when the install cannot be upgraded in place, the sha256 is missing ("check again first") or differs from the last check's, or a job is already running. The job is killed after 30 minutes even when it prints nothing. A version the check cannot compare is an error, never up_to_date.

GET /api/health adds "paused": true while paused; status stays "healthy".

Prompt-cache control plane (operator API)

The gateway exposes prompt-cache operator endpoints under /api/gateway/prompt_cache/*. Provider prompt-cache controls affect process-local or remote provider state and require an admin principal in hosted user-auth mode.

Core endpoints:

  • GET /api/gateway/prompt_cache/capabilities?provider=...&model=...
  • GET /api/gateway/prompt_cache/stats?provider=...&model=...
  • POST /api/gateway/prompt_cache/set
  • POST /api/gateway/prompt_cache/update
  • POST /api/gateway/prompt_cache/fork
  • POST /api/gateway/prompt_cache/clear
  • POST /api/gateway/prompt_cache/prepare_modules

Behavior:

  • These routes use the runtime's AbstractCore prompt-cache client contract rather than directly depending on provider-instance access.
  • In local mode they delegate to the in-process provider.
  • In remote/hybrid mode they follow whatever /acore/prompt_cache/* surface the configured AbstractCore server exposes.
  • All core prompt-cache responses include operation and capabilities, with structured unsupported/error cases (code="prompt_cache_unsupported" / code="prompt_cache_error" / code="prompt_cache_unavailable").
  • These endpoints remain provider/model controls, not a Gateway-owned CachedSession persistence system.

Session lifecycle endpoints:

  • GET /api/gateway/sessions/{session_id}/prompt_cache/status
  • POST /api/gateway/sessions/{session_id}/prompt_cache/prepare
  • POST /api/gateway/sessions/{session_id}/prompt_cache/rebuild
  • POST /api/gateway/sessions/{session_id}/prompt_cache/clear

These routes derive a deterministic bounded namespace/key from session_id, bundle_id, bundle_version, flow_id, provider, model, optional template_id, and version. The private hash also includes the authenticated principal scope, so two hosted users using the same session id/provider/model do not collide in a shared provider control plane; the returned identity remains portable app-level data and does not expose that private scope. These routes expose three honest modes:

  • unsupported: provider/model does not expose prompt-cache support; responses include supported=false, ok=false, and capabilities.
  • keyed: gateway returns a stable runtime_hint/prompt_cache_key for Runtime/Core injection, but does not claim module preparation occurred.
  • local_control_plane: gateway uses supported provider operations such as prepare_modules, fork, set, clear, and stats.

status is read-only. prepare accepts optional modules (system_prompt, workflow_instructions, tools, pinned_attachments) and returns either provider operation results or a key hint. rebuild is clear-plus-prepare for providers that expose clear controls.

These identity-derived endpoints only see caches whose keys the gateway derived. To enumerate or bulk-clear the caches the runtime actually minted for a session, use the recommended session prompt-cache enumeration lane.

Durable bloc exact-reuse endpoints:

  • POST /api/gateway/blocs/upsert_text
  • GET /api/gateway/blocs/record
  • GET /api/gateway/blocs
  • POST /api/gateway/blocs/delete
  • GET /api/gateway/blocs/kv/manifest
  • GET /api/gateway/blocs/kv/list
  • POST /api/gateway/blocs/kv/ensure
  • POST /api/gateway/blocs/kv/load
  • POST /api/gateway/blocs/kv/delete
  • POST /api/gateway/blocs/kv/prune

These routes are the primary app-facing durable prompt-cache path:

  • create or identify a durable text bloc;
  • ensure or load a KV artifact for a target local provider/model;
  • use the returned prompt_cache_binding in later Runtime-backed generation;
  • list/delete/prune artifacts without reaching into provider-private cache state.

They delegate through Runtime's public AbstractCore host facade rather than proxying Core directly. They are operator-style host controls, so the routes themselves are not ledgered run execution; the ledgered exact-reuse path is the later LLM_CALL.params.prompt_cache_binding used inside real Runtime runs.

Host-local prompt-cache export/import admin aliases:

  • GET /api/gateway/prompt_cache/saved
  • POST /api/gateway/prompt_cache/save
  • POST /api/gateway/prompt_cache/load

These remain explicitly local/operator-oriented:

  • the route paths are compatibility aliases, but the implementation delegates to Runtime's public host facade:
  • saved -> list_prompt_cache_exports(...)
  • save -> prompt_cache_export(...)
  • load -> prompt_cache_import(...)
  • local bundle/file runtimes store these exports under the Gateway data dir at prompt_cache_exports/
  • remote and hybrid runtimes return code=prompt_cache_local_only
  • response payloads follow Runtime's host-local export/import contract, including operation, local_only, artifact_*, capabilities, and provider_response

Accounts and activity

The Accounts page (admin): users and entities in one list.

Route Purpose
GET /admin/accounts {accounts: [{id, tenant_id, kind: user \| entity, role: admin \| user \| entity, own, email_address, mailbox{state: connected \| receive_only \| not_connected \| paused \| unavailable, address, provider, reason}, runtime_id, active, entity_state, actions{email, logs, workspace, preferences, rotate, manage, delete, suspend: {available, reason}}}]}, sorted admins, users, entities, then id; entities_warning when the entity list could not be read
PUT /admin/accounts/{id}/active {active} → the updated row. Users: false = deactivated (signed out, cannot sign in); 409 {message} for your own account ("You can't deactivate your own account.") or the last active admin. Entities: false = suspended (entity state paused, its door credential off, an open visit closed); true = resumed (the state it had before is restored, stored in <data_dir>/auth/entity_suspended.json)
GET /admin/accounts/{id}/activity ?limit=100&kind=sign_in,run,… (admin)
GET /me/activity the same for the signed-in account
GET /me/accounts any signed-in account: {accounts: [rows], scope: "own"}, the same row shape — your own row plus one row per entity you created; actions only an admin can take are unavailable with the reason
GET /me/accounts/{id}/activity the activity of your own account or of an entity you created; any other id answers 404

Who sees which account

An admin sees every user and every entity (GET /admin/accounts; non-admins get 403 there). Anyone else sees only themself and the entities they created. Entity rows carry created_by ({tenant_id, user_id} of the account whose POST /entities created it, or null).

This is enforced on every entity route, not only on the Accounts page:

  • POST /entities writes created_by into the new home's manifest.json. Entities created before this field existed have none and are visible to admins only: no creator is guessed and no existing manifest is rewritten.
  • GET /entities lists only the entities the caller may see.
  • Every /entities/{name}/… route (inspect, card, state, chat, visit, summon, workspace, tool policy, replay, …) checks first. An entity you may not see answers exactly like one that does not exist: 404 with the same sentence, so names cannot be probed this way.
  • POST /entities/meets/open needs both entities visible; /entities/meets/{id} answers 404 for a meet with an entity you may not see.
  • Creating an entity under a name another account already holds answers 409 ("That name is taken …"): entity names are unique per gateway, so this is the one place a name's existence shows. This holds across runtimes: an entity living in another user's runtime, or a user account with that name, also answers 409 (POST /entities/{name}/validate reports the same error).
  • Visibility is not management: entity writes that are admin-only (state, tool policy, prompt, substrate, …) stay admin-only for the entities you created.

Without user accounts (the single-operator gateway) every entity is visible, as before.

The email address and mailbox of a row come from the resolver GET /me/email uses. An action that cannot apply says why in reason: an entity has no mailbox ("Entities can't have their own mailbox yet: mailboxes belong to a user's runtime." — entity runs use their own runtime, without the host's mailbox), no token to rotate (its credential is discarded at creation) and no delete ("An entity's name is kept for life; suspend it instead.").

Activity answers {events: [{ts, kind: sign_in \| token \| run \| automation \| email \| account, title, detail, run_id, observer_path, ok, ts_local}], source: "audit_log", oldest_ts, truncated, note}, newest first. It reads <data_dir>/audit_log.jsonl and its rotated files backwards within a fixed read budget (truncated: older entries were not read). Events come from two explicit tables: the request lines (sign-ins, sign-out, runs started with their run_id, automation commands, account changes by an admin) and the typed email events (mailbox connected, tested, notification sent or not, …). The audit log records writes only: read-only requests (page views, token use on reads), mail received and what agents send with their email tools are not in it, and note says so. observer_path opens the event in the Observer app: /apps/observer/#run/<run_id> for a run, /apps/observer/#automations for an automation event; ts_local is ts in the gateway's local time (ISO 8601 with its offset).

A run id is created by the run-start route, so it is not in the request path: POST /runs/start and POST /runs/schedule write run {run_id, workflow, bundle_id, bundle_version, entrypoint, scheduled} on their audit line (workflow is the entrypoint's name, else the bundle id). A run event's detail is that workflow name. A run-start line written by a gateway older than 0.10.0 has no run id: its detail is "Run id not recorded (before this version)" and it has no run_id and no observer_path (nothing is guessed from times). A notification event's detail is its kind in words: "Approval needed", "Job failed", "Job finished", "Automation result", "Automation failed", "Test notification"; any other kind is shown as written.

Account preferences

The gateway keeps each account's client preferences, so every app and every device of that account see the same choice and no client keeps it. The gateway defines the defaults; an account may override them for itself.

Route Purpose
GET /api/gateway/accounts/{account}/preferences {ok, account: "tenant:user", can_edit, preferences: {default_workflow: {<interface>: value \| null}}, declared: {default_workflow: {label, help}}, apps: [row]}
PUT /api/gateway/accounts/{account}/preferences {"default_workflow": {<interface>: "bundle:flow" \| "catalog:bundle:flow" \| null}} → the GET answer

{account} is me (the caller), name, tenant:name or an entity's slug. Who: anyone for their own; an admin for any account; an entity's preferences by an admin or the entity's creator. Anyone else gets 403 with "Only an admin or the account itself can read or change its preferences." (an entity: "Only an admin or 's creator can read or change its preferences."). An account that does not exist answers 404. A gateway older than 0.13.1 has no such route (404 Not Found); the Assistant and AbstractCode then keep their own choice as before.

Only keys the gateway declares are accepted. Today there is one, default_workflow: one entry per app interface (abstractcode.agent.v1, abstractassistant.agent.v1). null follows the gateway's per-app default, which is the admin setting agents.default_workflow.<interface> (Workflows → Default workflow per app, configuration.md); it is never copied into the account. A chosen workflow is stored without a version, so it runs its latest published version. The PUT replaces the named interfaces and keeps the others. It is refused with 400 {"detail": {"reason": "preference_refused", "message": <sentence>, "key": <key>}} for an unknown key ("Unknown preference 'theme': this gateway declares default_workflow."), an interface that is not an app's, or a workflow the account may not run for that app ("…refused: 'helper@1.0.0:agent' declares abstractassistant.agent.v1, not abstractcode.agent.v1"). A refused PUT changes nothing. The change is audited on the target account's activity ("Preferences changed").

Each apps row: {interface, label, app, help, value, state, reason, gateway_default, gateway_default_label, effective, choices}.

Field Meaning
value the account's choice, null = the gateway default
state default (null), set (a choice that runs), broken (a choice that no longer runs; reason says why and what to do, e.g. "coder-two:agent no longer runs for alice: … Pick another workflow or Gateway default.")
gateway_default {available, name, value, workflow_id, reason}: what the admin's per-app default resolves to now
gateway_default_label "Gateway default ()", or "Gateway default (unavailable)"; every client shows it verbatim as the first option
effective what a run of this app starts for this account: {source: account \| gateway, available, name, value, workflow_id, bundle_id, bundle_version, flow_id, registry_scope, reason}
choices the workflows the account may run for this app: [{value, label, name, workflow_id, bundle_id, bundle_version, flow_id, registry_scope}], latest versions, never one an admin made unavailable to users or archived; label is shown verbatim

For me the choices are the caller's own (shared workflows available to them and their own workflows). For another account they are the gateway's shared workflows that account may run: an admin is never offered a user's private workflows.

Runs: an app starts the account's choice as an explicit workflow, or flow_id: "@default" with its interface when the value is null. @default keeps meaning the admin's per-app default.

Accounts rows (GET /admin/accounts, GET /me/accounts) carry actions.preferences {available, reason} (every user and entity that is not archived).

Email

Per-user email (email.md): every route acts on the caller's own account, resolved from the authenticated principal (never from a path or body id); entities are refused (403). No response ever carries a password, token or OAuth client secret.

The email address is the user's own address (sign-in codes, notifications, the default allowed recipient; no password); the mailbox is the connection the user makes so their agents and automations can read and send mail as them.

Your email address and mailbox (/api/gateway/me/..., every signed-in human):

Route Purpose
GET /me/email settings and status (mailbox{state, address, provider, reason} as in /admin/accounts): configured, address, imap, smtp, auth_kind, secret_set, policy, limits (+ usage), status{last_test, last_ok, last_error{code, cause, fix}}, watcher{state, last_poll, cursor, received}, admin_enabled, effective_enabled; for the account page: email_address (your email address as stored), registered_address ("self" for runs: that address, else the mailbox's own), email_available, notifications{job_failed, approval_needed} + notifications_unavailable_reason, agent_tools{on, available, unavailable_reason, active}, oauth_providers[{id, available, reason}]; send_capable (false for a mailbox stored without its SMTP leg — mailbox.state: receive_only)
POST /me/email/discover {address} → {address, domain, found, source, provider, imap, smtp, username, tried, defaults} (known providers, autoconfig, ISPDB, SRV, MX); defaults is what the form pre-fills: {imap{host, port, security}, smtp{…}, login, source: discovered \| standard, provider, message} (AbstractCore server_defaults: the discovered servers, else imap.<domain> 993 SSL and smtp.<domain> 465 SSL, "Standard settings for — change them if your provider uses others."); 400 for a non-address
PUT /me/email connect (save + test): {address, password, username?, display_name?, imap?{host, port, security, folder}, smtp?{host, port, security}, test}; username defaults to the discovered login, else the address; display_name to the stored name, else the address's local part; connecting sets your email address when it is empty (audited email.address_changed, reason mailbox_connected); without imap and smtp the servers are discovered (discovery{source, provider, tried} in the answer; none found = 400 email_discovery_failed with tried); signs in to both servers first (unless test: false), stores nothing on a failure and names the failing step (detail.step: imap | smtp, detail.message); a leg left out is filled with the domain's standard server (or the discovered one) and the test signs in to BOTH, so a connected mailbox can always send (an unreachable outgoing server is detail.step: smtp, nothing stored); tools_reloaded: the caller's workflows were reloaded so agents' toolsets follow
PUT /me/email/address {address} — your email address, stored on your user record ("" clears it; 400 email_invalid_settings for anything but one plain address)
PUT /me/email/notifications {job_failed?, approval_needed?} — the two notification switches (both on by default); the earlier {email: {...}} body is accepted and mapped; answers like GET /me/email
POST /me/email/test per-leg result {imap, smtp, ok, message}; message: "Test passed: signed in to imap.x and smtp.x." or the failing step and its cause ("Sign-in refused by imap.x — check the password. (…)")
DELETE /me/email disconnect: credentials and cursor deleted (policy and limits kept); tools_reloaded
PUT /me/email/policy {mode: "allowlist" \| "denylist", always_allow?: [address \| domain], always_deny?: [address \| domain]}; a list given replaces that list. Precedence: own address → allowed; Always denied → refused; Always allowed → allowed; else the mode (allowlist = Only the Allowed list → refused, denylist = Anyone not on the Denied list → allowed). A domain covers its subdomains. The older {mode, entries} is accepted (entries = the mode's list)
POST /me/email/policy/check {to?, cc?, bcc?, addresses?} (addresses = To) → per-recipient verdicts with source (self, always_deny, always_allow, mode)
PUT /me/email/limits {per_hour, per_day}
PUT /me/email/folder {folder} — the folder your mailbox is read from (empty = INBOX); the connection is kept and nothing is tested; 404 email_not_configured without a mailbox
PUT /me/email/enabled {enabled} — "Use this mailbox" (off keeps the settings; stops watching, sending and notifications); tools_reloaded
PUT /me/email/agent-tools {enabled} — your agents' email tools (default off; 409 email_disabled while "Agent email tools for users" is off for you; active only with a connected, allowed mailbox); reloads your workflows so toolsets follow (tools_reloaded)
GET /me/email/oauth/clients which providers have a gateway OAuth client (no secrets)
POST /me/email/oauth/start {address, provider: google \| microsoft, client_id?, client_secret?, tenant?, flow?: device \| loopback} → device code (user_code, verification_uri) or authorization_url. token_endpoint, authorization_endpoint, device_authorization_endpoint and scopes are accepted only from an administrator with provider: "custom" (own client id); otherwise 403 email_oauth_override_refused
POST /me/email/oauth/poll {flow_id} → {pending: true} or the connected account
POST /me/email/oauth/finish {flow_id, wait_s} — waits up to 60 s for the approval; tools_reloaded once connected
POST /me/email/oauth/cancel {flow_id}
GET /me/notifications events and choices, channel availability, outbox summary
PUT /me/notifications {email: {job_failed, approval_needed: bool}}; the earlier kinds are accepted (automation_failed counts for job_failed; automation_result and job_finished are per-automation / per-run options now)
POST /me/notifications/test sends one test notification now → {ok, sent, reason_code, message, limit, state, error?}: reason_code is null (sent) or no_mailbox, mailbox_paused, rate_limited, queued_behind (+ queued_behind: how many), send_failed; message is the sentence to show ("Sent to x@y.", "Not sent: hourly limit reached (20 of 20 this hour) — resets at 14:05.", in the gateway's local time); limit {window: hour \| day, limit, used, resets_at} when a send limit held it back

Administrators (status and the switch only; administrators never read mail):

Route Purpose
GET /admin/users every row carries email_address and mailbox{state: connected \| receive_only \| not_connected \| paused \| unavailable, address, provider, reason} from the same resolver as GET /me/email (so your own row matches your card; email stays the raw record field); each human row also carries email_account: {configured, address, state, admin_enabled, agent_tools_available, capabilities} (capabilities: {email, email_agent_tools} as {value, source: user \| gateway \| built-in}; user = a per-user override)
GET /admin/users/{user_id}/email configured, address, auth_kind, user_enabled, admin_enabled, effective_enabled, status (last test / last error), capabilities ({value, source: user \| gateway \| built-in}), agent_tools (available, user_enabled, active), watcher, state
PUT /admin/users/{user_id}/email {enabled?, agent_tools?, inherit?: ["email", "email_agent_tools"]} — per-user capabilities; enabled: false = no watcher, no sending, no notifications (settings kept); agent_tools = Agent email tools available
GET /admin/email/capabilities capabilities[{id, label, description, per_user, advanced, default, built_in_default}]: email "Mailboxes for users" (on), email_agent_tools "Agent email tools for users" (on) and email_recovery "Sign-in by email" (on)
PUT /admin/email/capabilities {email?, email_agent_tools?, email_recovery?, reset?: [...]}; tools_reloaded: how many built hosts were rebuilt so every user's agents follow the change
GET /admin/email/oauth-clients bring-your-own OAuth clients: client_id, client_secret_set, tenant per provider
PUT /admin/email/oauth-clients/{provider} {client_id, client_secret?, tenant?}; an empty client_id removes the provider's client; omitting client_secret keeps the stored one for the same id

Sign-in page (public):

Route Purpose
GET /session/recovery {available} — true when sign-in by email is on (email_recovery, default on) and at least one account of this gateway has email
POST /session/recovery/request {user_id, tenant_id?, purpose?: sign_in (default) \| reset_token} → {sent: true, to: "l•••@•••", expires_in_s, message}, {sent: false, reason_code: "no_email_address", message} (also for an unknown account), {sent: false, reason_code: "no_mailbox", message} (an address but no mailbox to send from), {sent: false, reason_code: "send_failed", message} (the mail server refused or could not be reached; the request waits up to 12 s for the real outcome) or {sent: false, reason_code: "too_many_requests", retry_after_s, message}; 404 recovery_off when sign-in by email is off (email.md)
POST /session/recovery/redeem {user_id, tenant_id?, purpose?, code, remember?} → a browser session; reset_token also returns the new token once. A wrong, expired or used code answers 401 recovery_code_refused

Errors carry {"detail": {"reason_code", "message", "cause", "fix", "retryable"}}: 400 invalid settings or a policy refusal (email_policy_refused), 403 entity or a refused OAuth override (email_oauth_override_refused), 404 no account (email_not_configured), 409 turned off (email_disabled) or credentials missing, 422 the mail server refused (email_auth_failed, email_tls_failed, email_unreachable, …), 429 send limit (email_rate_limited).

curl -sS -H "$AUTH" "$BASE_URL/api/gateway/me/email"
curl -sS -X POST -H "$AUTH" "$BASE_URL/api/gateway/me/email/test"
curl -sS -X POST -H "$AUTH" "$BASE_URL/api/gateway/me/notifications/test"

Runs: _runtime.notify = {"on": ["finished", "failed"], "channels": ["email"]} in a run's input asks for a notice when it ends (unknown values are dropped). _runtime.email_account and _runtime.email_allowed_recipients are set by the gateway; client-supplied values are removed.

Deprecated aliases (admin only, the calling administrator's own account; removed in a later minor release): GET /email/accounts, GET /email/messages (mailbox, since, status, limit), GET /email/messages/{uid} (whole body, content_trust: "untrusted"), POST /email/send ({to, cc?, bcc?, subject, body_text?, body_html?}; the recipient policy and limits apply).

Troubleshooting and common questions: faq.md.