AbstractGateway — API overview¶
The HTTP API is implemented with FastAPI under the /api prefix:
- Health: GET /api/health (unauthenticated; status, runner, watchdog: {enabled, limit_s, last_tick_age_s} — see troubleshooting.md). The previous process's last watchdog incident is admin-only: GET /api/gateway/host/runner → last_hang: {at, blocked_s, reason, top_frame, dump_path, file, line} or null
- Gateway surface: /api/gateway/* (durable runs + operator tooling)
The API is documented at runtime:
- OpenAPI JSON: GET /openapi.json
- Swagger UI: GET /docs (use Authorize to paste the bearer token)
Context: - In the AbstractFramework ecosystem, UIs and automations call this API to operate AbstractRuntime runs. - Architecture diagram and core concepts: architecture.md
Route families¶
This page covers the run contract, artifacts, discovery, media, models and host state. Other route families are documented next to the feature they serve:
| Routes | Purpose | Reference |
|---|---|---|
/api/gateway/session/login, /session/logout, /session/claim, /me |
browser sessions, one-time sign-in links, the current principal | security.md, first-run.md |
/api/gateway/admin/users, /admin/runtime-reservations |
user accounts and retained runtimes (admin) | security.md |
/api/gateway/admin/accounts, /admin/accounts/{id}/active, /admin/accounts/{id}/activity, /me/accounts, /me/accounts/{id}/activity, /me/activity |
the Accounts page: users and entities in one list (admin), your own account and your entities (everyone), the Active switch, activity from the audit log | below |
/api/gateway/admin/runtime-config |
runtime settings (admin) | configuration.md |
/api/gateway/accounts/{me\|account}/preferences |
one account's client preferences: the default workflow per app (the account itself, an admin, an entity's creator) | below |
/api/gateway/workspace/policy, /workspace/policy/{account}, /sessions/{id}/workspaces, /workspace/effective/{account} (GET and dry-run POST), POST /workspace/path-check |
workspaces, three levels: the gateway policy = the eligible set (read: everyone; write: admin), one account's default subset (admin, the account itself, an entity's creator; me = the caller), one conversation's subset (its owner, or an admin), the effective set the gateway enforces, and the path check run before saving a row ({path} → {path, normalized, absolute, exists, is_dir, valid, sentence}; any signed-in principal) |
below, security.md |
/api/gateway/admin/runtimes, ?account=<id>[&tenant_id=<t>] |
the runtime inventory (admin); account keeps only the planes that account owns and echoes filter |
console.md |
/api/gateway/network, /network/restart |
network exposure, addresses, reverse proxy | configuration.md |
/api/gateway/apps/*, /apps/handover/{code}, /apps/tui-handover, /api/gateway/apps/desktop-handover |
browser apps, terminal apps, the Assistant and its sign-in | apps.md |
/api/gateway/runs/{run_id}/workspace, /workspace/files, /workspace/content |
browse and preview a run's folder (the run's owner) | below |
/api/gateway/skills, /admin/skills/reseed |
the skills shelf | below, configuration.md |
/api/gateway/about |
versions this gateway runs (no sign-in) | below |
/api/gateway/engines/* |
local engine installs | engines.md |
/api/gateway/models/download*, /models/downloads* |
model download jobs and their event stream | model-downloads.md |
/api/gateway/host/* |
host state, pause, restart, update, tray | Host state, Host control |
/api/gateway/backlog/*, /reports/*, /triage/*, /processes |
operator tooling | maintenance.md |
/api/gateway/entities/* |
summoned entities | entities.md |
/api/gateway/automations*, /api/gateway/trigger-sources |
automations: recurring workflows, their occurrences and attention | automations.md |
OpenAI API¶
The OpenAI-compatible API is served at /v1 on this Gateway listener
(/v1/models, /v1/chat/completions, /v1/embeddings, ...); callers use their
gateway token as the API key. /core/v1 answers 308 to /v1 (deprecated).
Status and logs: GET /api/gateway/openai-api[/logs]; admin changes:
/api/gateway/admin/core-endpoint. See openai-api.md for the
supported surface, access settings and errors.
Auth¶
By default, /api/gateway/* is protected by GatewaySecurityMiddleware (bearer token + origin allowlist).
See: security.md.
All examples below assume:
export BASE_URL="http://127.0.0.1:8080"
export AUTH="Authorization: Bearer $(cat "$ABSTRACTGATEWAY_DATA_DIR/auth/bootstrap-admin-token")"
Provider connections¶
Gateway-owned provider connections let users create reusable cloud, local, or
OpenAI-compatible endpoints without putting raw API keys in workflow JSON or
browser storage. The API route is named provider-endpoint-profiles; the
console presents them as provider connections.
GET /api/gateway/config/provider-endpoint-profiles: list visible profiles.POST /api/gateway/config/provider-endpoint-profiles: create a user- or admin-owned profile.POST /api/gateway/config/provider-endpoint-profiles/discover-models: discover models for a draft or saved profile by calling the configured provider family and base URL with the entered or server-side key. The raw key is never returned.PUTorDELETE /api/gateway/config/provider-endpoint-profiles/{profile_id}: update or delete a profile.
Enabled profiles appear in GET /api/gateway/discovery/providers as virtual
providers such as endpoint:office-vllm. Model discovery through
GET /api/gateway/discovery/providers/{provider_name}/models returns either the
fixed profile allowlist or the live endpoint model catalog.
Core workflow lifecycle¶
1) List bundles (bundle mode)¶
curl -sS -H "$AUTH" "$BASE_URL/api/gateway/bundles"
Each item also says where the bundle came from and what it does:
source:shipped(a file the gateway package ships),published(published from AbstractFlow through this gateway: the publish route's stamp in the manifest metadata) orimported(anything else: an uploaded.flow, or a file copied into the folder);description: the default entrypoint's description (""when it has none).
Upload a bundle:
curl -sS -H "$AUTH" \
-F "file=@./my-bundle@0.1.0.flow" \
-F "overwrite=false" \
-F "reload=true" \
"$BASE_URL/api/gateway/bundles/upload"
2) Start a run¶
curl -sS -H "$AUTH" -H "Content-Type: application/json" \
-d '{"bundle_id":"my-bundle","input_data":{"prompt":"Hello"}}' \
"$BASE_URL/api/gateway/runs/start"
If you need a specific entrypoint:
curl -sS -H "$AUTH" -H "Content-Type: application/json" \
-d '{"bundle_id":"my-bundle","flow_id":"ac-echo","input_data":{"prompt":"Hello"}}' \
"$BASE_URL/api/gateway/runs/start"
Every start answers {run_id, runner_warning, resolved_workflow}.
resolved_workflow names the workflow the run really runs:
{workflow_id, bundle_id, bundle_version, flow_id, registry_scope, name,
source: "gateway_default" | "client", interface}. The same object is kept in
the run's inputs as input_data.workflow_selection (written by the gateway;
a value sent by the client is replaced), so GET /runs/{run_id}/input_data
tells a restored conversation how its workflow was chosen.
Evidence: request/response models live in src/abstractgateway/routes/gateway.py (StartRunRequest, start_run).
The gateway default agent workflow (flow_id: "@default")¶
An agent client (AbstractCode, the Assistant, the Telegram bridge) can let the gateway choose the workflow for an agent interface:
curl -sS -H "$AUTH" -H "Content-Type: application/json" \
-d '{"flow_id":"@default","interface":"abstractcode.agent.v1","input_data":{"prompt":"Hello"}}' \
"$BASE_URL/api/gateway/runs/start"
interfaceis required with@default(400 without it);bundle_idandbundle_versionare not accepted with it.- The default is resolved at every start, so a change applies to the next
new turn of any conversation that sends
@default. - When the default cannot run (its workflow is gone, deprecated, or does not
declare the interface), the start is refused with 409 and a message naming
the setting (
agents.default_workflow.<interface>) and where its value comes from. The gateway never quietly runs another workflow instead. POST /runs/scheduleaccepts the sameflow_id: "@default"+interface(the schedule then targets the version resolved at that moment).
What the default is for each interface (readable without admin rights):
GET /bundles and GET /workflow-catalog carry
"default_agent_workflows": {
"abstractcode.agent.v1": {"workflow_id": "basic-agent@0.0.5:81795ea9", "bundle_id": "basic-agent",
"bundle_version": "0.0.5", "flow_id": "81795ea9", "registry_scope": "private",
"name": "basic-agent", "source": "default"}
},
"default_agent_workflows_unavailable": {
"abstractassistant.agent.v1": {"source": "default", "value": null,
"reason": "no host workflow declares abstractassistant.agent.v1; the Assistant uses its built-in orchestrator"}
}
and every entrypoint row carries is_agent_default and
agent_default_interfaces. These are different from default_bundle_id
(the bundle a bare flow_id falls back to), from a catalog record's
is_default (its default version) and from a bundle's default_entrypoint.
The setting itself is agents.default_workflow.<interface> (see
Configuration). In GET /admin/runtime-config, each
agents.default_workflow.<interface> row also carries the plain words the consoles show, from one
table in agent_defaults.py (INTERFACE_TABLE):
| Field | Meaning |
|---|---|
interface, label, app, help |
the interface id, its plain name ("AbstractCode — chat agent"), the app that asks for it (null when none does) and one sentence of help |
group |
apps (an app asks the gateway for it as its default agent) or other (declared by a workflow; no app asks for it by default) |
state |
builtin (nothing saved; the gateway default runs: the shipped bundle, else the newest available workflow), none (nothing saved and no workflow declares the interface; reason says so), set (a saved value that runs) or broken (a saved value that no longer resolves) |
value |
the saved value, else the gateway default, else null |
reason |
for broken: "Broken: workflow bundle 'coding-agent' is not on this gateway — pick another workflow or the gateway default."; for none: "no workflow on this gateway declares |
For VisualFlow bundles, Gateway runs the packed JSON through AbstractRuntime.
Structured LLM/Agent schemas are Runtime/Core-owned: response remains textual,
and schema-conformant object values are available through the node data output
for data edges such as Break Object and Switch.
Durable session replay (use_session_history)¶
Thin clients do not need to carry conversation transcripts. Passing
"input_data": {"use_session_history": true} together with a session_id
makes the gateway seed the run's context.messages from the session's prior
COMPLETED root runs before the run starts: the run store is the durable
transcript.
curl -sS -H "$AUTH" -H "Content-Type: application/json" \
-d '{"bundle_id":"my-bundle","session_id":"sess-1","input_data":{"prompt":"and what did I say before?","use_session_history":true}}' \
"$BASE_URL/api/gateway/runs/start"
Rules (the model-vs-display divergence contract — what the model replays is deliberately narrower than what history views display):
- Client-provided non-empty
context.messagesalways win; the seed never overwrites them. An EMPTY clientcontext.messageslist does not count as a transcript — the seed still runs (leave outuse_session_historyto start without history). - Only COMPLETED root runs of the session contribute, as strictly alternating user/assistant pairs. FAILED and CANCELLED turns are invisible to replay by design (a promptless answer or answerless prompt would seed a dangling message and invite re-answering a stale ask); history views still show them.
- Steering/operator guidance injected mid-run is not replayed.
- The history window: replay keeps the most recent turns that fit 50,000
estimated tokens (AbstractRuntime
HISTORY_REPLAY_MAX_TOKENS), as whole messages, newest first. No message is cut, and there is no message-count or character cap. The model can use the rest of its context window. When older turns are dropped, the oldest replayed message starts with a labeled[#TRUNCATION: ...]line. A newest turn that alone exceeds 50,000 tokens is kept whole (oversize_turn_kept: true). - The retired caps
input_data.session_history_max_messagesandinput_data.session_history_max_charsare ignored and listed in_runtime.session_history.ignored_inputs. The environment variablesABSTRACTGATEWAY_SESSION_HISTORY_MAX_MESSAGESandABSTRACTGATEWAY_SESSION_HISTORY_MAX_CHARSare not read. - Failures degrade to a labeled
_runtime.session_history#FALLBACKnote and an unseeded start — never a blocked run. Success records the window on the run as_runtime.session_history:{seeded, policy, max_tokens, token_estimator, replayed_messages, replayed_tokens, dropped_messages, dropped_tokens, dropped_counts_complete, oversize_turn_kept}.dropped_counts_complete: falsemeans replay stopped reading once the window was full: older turns were dropped too, and they are not counted. - Entity lanes never ride this: their transcript authority is the entity home
(
_visit.history/ the chat driver), not the run store.
Evidence: _seed_session_history in src/abstractgateway/hosts/bundle_host.py
and abstractruntime.session_history.session_chat_messages.
2b) Schedule a run (bundle mode)¶
POST /api/gateway/runs/schedule starts a scheduled parent run that launches the target workflow as child runs over time.
Example (run 3 times, every hour, starting now):
curl -sS -H "$AUTH" -H "Content-Type: application/json" \
-d '{"bundle_id":"my-bundle","flow_id":"ac-echo","input_data":{"prompt":"Ping"},"start_at":"now","interval":"1h","repeat_count":3,"share_context":true,"session_id":"sess-1"}' \
"$BASE_URL/api/gateway/runs/schedule"
Notes:
- start_at: ISO 8601 timestamp (recommended) or "now".
- interval: e.g. "15m", "1h", "2d". If omitted, runs once.
- repeat_count: if omitted and interval is set, repeats forever. Alternatively use repeat_until (ISO 8601).
- To stop a schedule, cancel the scheduled parent run via POST /api/gateway/commands with type cancel.
Evidence: ScheduleRunRequest, start_scheduled_run in src/abstractgateway/routes/gateway.py.
2c) Shared workflow catalog¶
Private /api/gateway/bundles routes are scoped to the signed-in user's routed
runtime, and you may change the registry you own. The gateway's own bundle
directory is shared by every user, so writing it — upload, delete, reload,
deprecate, and POST /visualflows/{flow_id}/publish — requires an admin
principal and otherwise returns 403. Listing and running are unaffected. See
security.md for the full rule.
Shared/default workflows use the Gateway workflow catalog instead:
curl -sS -H "$AUTH" "$BASE_URL/api/gateway/workflow-catalog"
Admin-only catalog operations live under
/api/gateway/admin/workflow-catalog/*:
- upload or promote immutable
.flowversions; - move a bundle's default pointer;
- set ACLs;
- deprecate, block, or tombstone a version without deleting bundle bytes.
Start a catalog workflow in the requesting user's runtime by setting
registry_scope:
curl -sS -H "$AUTH" -H "Content-Type: application/json" \
-d '{"registry_scope":"tenant_catalog","bundle_id":"basic-agent","flow_id":"root","input_data":{"prompt":"Hello"}}' \
"$BASE_URL/api/gateway/runs/start"
If bundle_version is omitted, Gateway uses the admin-managed catalog default
pointer. Exact older versions keep working until that specific version is
deprecated, blocked, or tombstoned.
Catalog scope is explicit: omitting registry_scope starts only private
runtime bundles. Flow/schema inspection for catalog workflows should use the
ACL-aware catalog endpoints:
GET /api/gateway/workflow-catalog/{bundle_id}/versions/{bundle_version}/flows/{flow_id}GET /api/gateway/workflow-catalog/{bundle_id}/versions/{bundle_version}/flows/{flow_id}/input_schema
framework_catalog is reserved but not loadable yet; use tenant_catalog.
2d) Docs Q&A (docs-qa catalog bundle)¶
docs-qa is the shared transport for docs-grounded assistant panels (the
unified top-bar drawers). The contract: the CALLER supplies its own corpus
(typically its llms.txt text) — the bundle never guesses one, so answers are
never silently grounded on another app's docs.
Fresh installs need no manual publish: the gateway ships docs-qa in the
wheel and boot idempotently publishes it into the tenant catalog
(publish-if-absent by exact version; an admin's default pointer, tombstones,
and publisher attribution are never touched; publisher system:gateway-boot).
The publish is skipped on custom-bundle deployments whose private registry
carries no LLM-bearing flow (it would add a boot requirement they never had)
and can be disabled with ABSTRACTGATEWAY_AUTO_PUBLISH_SHIPPED=0 — the
manual upload below then remains the path.
curl -sS -H "$AUTH" -H "Content-Type: application/json" -d '{
"registry_scope": "tenant_catalog",
"bundle_id": "docs-qa",
"bundle_version": "0.1.1",
"flow_id": "docsqa001",
"session_id": "myapp-docs-assistant:<one id per conversation>",
"input_data": {
"prompt": "How do I publish a workflow bundle?",
"docs": "<your llms.txt text>",
"app": "MyApp",
"use_session_history": true
}
}' "$BASE_URL/api/gateway/runs/start"
Conversation history comes from the run's session, never from the caller
(0.1.1): start every question of one conversation with the same session_id
and use_session_history: true. The gateway replays the session's earlier
turns through the runtime's history window (the newest whole turns up to
50,000 tokens) and the bundle's LLM call includes them; a new conversation is
a new session_id. 0.1.0's question + history inputs (last 12 messages
kept) are gone.
Then poll GET /runs/{run_id} (or stream the ledger); the answer is
output.response, and session_history is the window's receipt
(replayed_messages, dropped_messages, ...) to show when earlier messages
were not replayed. provider/model/temperature may ride input_data to
override gateway defaults. Answers cite section headings and say plainly when
the docs do not answer — the bundle refuses to invent endpoints or behavior.
Docs Q&A must never route through entity chat (a visit is billable and forms
memories).
Run-level skills selection¶
input_data.skills (a list of skill NAMES) attaches curated skills to any
run started through /runs/start:
curl -sS -H "$AUTH" -H "Content-Type: application/json" -d '{
"bundle_id": "basic-agent",
"input_data": {"prompt": "…", "skills": ["agora-collaboration"]}
}' "$BASE_URL/api/gateway/runs/start"
Trust semantics (the same abstractskill gate as GET /skills and the
workforce spawn lane — one gate, never a second resolver): VALIDATED skills
activate and their index lands in the run's _runtime.skills_block
(byte-stable for the whole run) with the read_skill tool made reachable;
UNVERIFIED skills are held; advisory-BLOCKED skills never ride. Every
outcome is recorded as a labeled verdict in _runtime.skills_resolution
(requested/active/verdicts/resolved_tree_hashes) — nothing is
silently dropped. Agent-node subruns inherit the block verbatim with
read_skill appended to explicit child allowlists (empty allowlists keep
registry defaults). A caller-supplied _runtime.skills_block is never
overwritten; the selection is then ignored with a labeled verdict.
The gateway serves the corpus every Docs assistant grounds on (the kit's DocsAssistantDrawer in the console and the apps):
curl -sS -H "$AUTH" "$BASE_URL/api/gateway/docs/corpus" # the gateway's own llms.txt
curl -sS -H "$AUTH" "$BASE_URL/api/gateway/docs/corpus?app=code" # an app's llms.txt
Returns {app, source, chars, text}. With app=<id> (code, flow, observer, continuum, entity) the text is the llms.txt the running app serves from its own build (GET /llms.txt on its loopback port, text/plain only; source app:<id>:llms.txt); an app that is not running, an unknown id, or an app that does not serve one is a 404 that says which. Without app (or app=gateway): Resolution order: the
ABSTRACTGATEWAY_DOCS_CORPUS env override first (set-but-missing is an honest
404 naming the checked candidates, never a silent fallback), then the repo
llms.txt in dev checkouts, then the corpus packaged with the wheel.
Skills shelf¶
GET /api/gateway/skills lists the skills of the gateway's shelf:
{skills: [...], shelf, shelf_source, bundled_version, warnings}.
shelf_source says where the shelf comes from: stored (the saved setting
skills.shelf), env (a legacy launch environment value), seeded (the
gateway's own copy in <data dir>/skills/registry, kept up to date from the
curated shelf that ships with AbstractSkill at each start), checkout (a
framework checkout, used only when the gateway's own copy is missing) or
none. An empty list always comes with a warning that says why and what to
do; warnings are plain sentences meant to be shown as they are.
POST /api/gateway/admin/skills/reseed (admin) refreshes the gateway's own
copy now and answers the seed report (added, updated, unchanged, the
kept_* lists with what was kept and why, bundled_version,
previous_version). An edit made in that folder is never overwritten.
3) Replay the ledger (cursor-based)¶
Ledger pages are replayed using after as “number of items already consumed”.
curl -sS -H "$AUTH" "$BASE_URL/api/gateway/runs/<run_id>/ledger?after=0&limit=200"
Response shape:
- items: list of durable ledger records
- next_after: the next cursor to use
Evidence: src/abstractgateway/routes/gateway.py (get_ledger).
3b) Replay ledgers for multiple runs (batch)¶
Use POST /api/gateway/runs/ledger/batch to reduce request fanout when observing many runs/subflows.
curl -sS -H "$AUTH" -H "Content-Type: application/json" \
-d '{"limit":200,"runs":[{"run_id":"<run_id_1>","after":0},{"run_id":"<run_id_2>","after":0}]}' \
"$BASE_URL/api/gateway/runs/ledger/batch"
Evidence: src/abstractgateway/routes/gateway.py (get_ledger_batch).
4) Stream ledger updates (SSE)¶
SSE is an optimization; clients should always be able to reconnect by replaying from the last next_after.
curl -N -H "$AUTH" "$BASE_URL/api/gateway/runs/<run_id>/ledger/stream?after=0"
Each ledger record arrives as event: step with an id: line (the records
consumed so far); a reconnect sends Last-Event-ID (or ?after=) and
resumes there. event: done closes the stream once the run is finished and
everything was sent.
Evidence: src/abstractgateway/routes/gateway.py (stream_ledger).
4b) Live replies: token deltas on the same stream¶
A run started with input_data._runtime.stream: true also sends the model's
reply while it is being written, on the same SSE stream:
event: llm.delta
data: {"kind":"llm.delta","run_id":"…","root_run_id":"…","parent_run_id":null,"node_id":"llm","call_id":"<step id>","seq":3,"text":" a natural","channel":"content","snapshot":false}
event: llm.delta_end
data: {"kind":"llm.delta_end","run_id":"…","root_run_id":"…","parent_run_id":null,"node_id":"llm","call_id":"<step id>","seq":24,"reason":"completed","snapshot":false}
call_idis the step id of the LLM call; the durablellm_callrecord with the samestep_idcarries the final answer and replaces the live text.channeliscontent(the answer) orreasoning(thinking text, already separated on the server; clients do not parse<think>themselves).seqcounts the events of one call (deltas and end).- Delta frames have NO
id:line: they are not ledger records, andLast-Event-IDnever moves because of them. - Snapshots. When a client connects (or reconnects) while a call is still
being written, it first receives one frame per open call and channel with
snapshot: trueand the text so far, then the live frames. A client drops its live bubbles on every (re)connect and applies the snapshots. - Order. A call's
llm.delta_endis sent after its durable record; the last short batch of text can arrive just after the record, so a client ignores deltas of a call whose record it already holds.donecomes after every live frame. - Sub-runs. A stream of a run also carries the deltas of every run below
it (a coding agent's delegated calls); a stream of a child run carries that
child's subtree only.
run_idnames the run that wrote the text,root_run_idits root. - Endings.
reasoniscompleted,failed,cancelledorunavailable.unavailablecomes with adetailsaying why that call did not stream (structured_output,remote_core,provider_cannot_stream,usage_unavailable,prompt_cache_unavailable,node_stream_off,sink_error); the answer is complete either way. When a run ends (stopped, failed) while a call is still open, the gateway closes that call withreason: "cancelled"or"failed"andsynthetic: true. - Nothing is capped: a call's text is kept whole until it ends, and a slow client catches up from it.
- Another user's run answers 404, and no delta of it ever reaches anyone else.
The switch is input_data._runtime.stream (true or false; anything else
is refused with 400). When a POST /runs/start request does not say, the
gateway setting agents.streaming_default decides (off unless saved); it is
never applied to POST /runs/schedule, the bridges or the entity loop. A
flow node whose LLM call sets stream: false is never streamed (its end says
detail: "node_stream_off"). GET /api/gateway/discovery/capabilities
advertises the feature to every client:
"streaming": {"deltas": true, "default": false, "run_field": "_runtime.stream",
"events": ["llm.delta", "llm.delta_end"], "subtree": true,
"endpoint": "/api/gateway/runs/{run_id}/ledger/stream",
"end_reasons": ["completed", "failed", "cancelled", "unavailable"]}
With serve --no-runner plus abstractgateway runner, the runner writes a
run's deltas to <data dir>/live/<root run id>.deltas.jsonl (readable by the
gateway's user only) and the API process reads them from there; the file is
deleted when the run ends (however it ends: completed, failed, stopped, or
stopped by the kill switch), and files of finished runs are removed at start.
Disk use: the file holds every delta of every call of the run and of its
sub-runs, and grows until the ROOT run ends; nothing is capped, so a long agent
run that makes many calls (a coding agent working for an hour) can leave a file
of several megabytes in <data dir>/live while it runs.
Evidence: src/abstractgateway/live_deltas.py, src/abstractgateway/routes/gateway.py
(stream_ledger, _apply_run_stream_switch, discovery_capabilities),
src/abstractgateway/hosts/bundle_host.py (install_live_delta_sink).
A run's workspace folder (browse and preview)¶
A run works in a folder on the gateway computer: the conversation's own
folder the gateway made (<data dir>/workspaces/session-…), or the folder the
client was started from. A run started without workspace_root gets the
conversation's folder (a run without a session gets its own
<data dir>/workspaces/<run>), and its file tools resolve relative paths
there, never in another workspace. This private workspace is always read &
write for its run and is never listed. The agent's workspace context lists the
run's allowed workspaces with their paths and modes
(security.md). Three routes let the person who started the run see
it; another user's run id answers 404.
GET /api/gateway/runs/{run_id}/workspace:
{"run_id": "…", "workspace_root": "/Users/me/Library/Application Support/abstractgateway/workspaces/session-chat-1-3f2a…",
"kind": "session", "session_id": "chat-1", "exists": true,
"host": {"hostname": "studio.local", "caller_is_this_machine": true},
"open_supported": true}
kind is session, run or launch_folder. open_supported is true only
for an admin sitting at the gateway computer, the only caller for whom
POST /runs/{run_id}/workspace/open can open the folder; elsewhere show the
path and the host name.
GET /api/gateway/runs/{run_id}/workspace/files?path=<folder>&recursive=false&limit=<n>:
{"path": "src", "entries": [{"name": "a.py", "path": "src/a.py", "type": "file", "size_bytes": 11, "mtime": "2026-09-25T10:00:00Z"}],
"truncated": false, "limit": null, "recursive": false,
"hidden": {"outside_links": 0, "blocked": 0, "other": 0}}
Folders come first, then files, by name. limit is optional; when the listing
stops there, truncated is true. hidden counts what is not shown: links
that lead outside the folder, and entries the workspace deny list blocks.
GET /api/gateway/runs/{run_id}/workspace/content?path=<file> streams the
whole file with its content type, Content-Disposition: inline,
X-Content-Type-Options: nosniff and Content-Security-Policy: sandbox, and
honours Range (206; 416 outside the file).
Paths are relative to the folder: absolute paths and .. are refused (400),
a link that leads outside is refused (403), the gateway's own marker file is
never listed nor served. The built-in deny list (credential folders such as
~/.ssh, and the gateway's data folder; see
Configuration) is never
listed, and reading inside it answers 404, even when the run's folder
contains it. The run's effective workspaces are applied again at every call, and nothing
else in the gateway's data folder is ever served. A run cannot be started
with a workspace_root inside the gateway's data folder either, except the
conversation folder the gateway made for the same user.
Workspaces¶
Three levels, the same shape at each: a posture (allowed_only = "Deny
everything, allow listed workspaces", any_except_denied = "Allow everything,
refuse listed workspaces"), a default_mode (ro | rw, the mode of every
unlisted directory under "Allow everything…") and rows {path, mode: "ro" |
"rw" | "deny"} (Read-only, Read & write, Refused). There is no shared
workspace. Full model and enforcement:
security.md.
Refusals. Every workspace write, the dry run and every run start refuse with HTTP 400 and one shape:
{"detail": {"reason": "workspace_refused", "message": "The gateway allows this workspace read-only: /archive.", "path": "/archive"}}
path names the offending path, or is null. Clients show message (plus
"Not saved." after a write). A permission refusal is a 403 whose detail is
a sentence.
Gateway: the eligible set (admin)¶
GET /api/gateway/workspace/policy (any signed-in principal) → {ok, policy}
with policy = {posture, default_mode, folders: [{path, mode}], builtin_refused[]
(read-only), max_attachment_bytes (read-only), summary}. Each row's mode is
the CAP: no level below may reach outside the set or raise a cap. A fresh
gateway answers any_except_denied, rw, no rows. The built-in refusals (the
gateway's data folder, credential folders) always apply; a non-admin gets
builtin_refused: [] and builtin_refused_hidden: true (those paths name the
gateway's home and data folder).
PUT /api/gateway/workspace/policy (admin): any subset of {posture,
default_mode, folders}. Named fields replace and the others are kept, so one
control is one request. Every path is checked like POST
/workspace/path-check (absolute, existing). Every row states its mode. These
are refused:
- a duplicate row;
- a read & write or read-only row inside a built-in refusal (the data folder,
except a conversation folder
<data>/workspaces/<name>or<data>/users/<tenant>/<runtime>/runtime/workspaces/<name>, and the credential directories):Workspaces: '<path>' is inside the built-in refused workspace '<built-in>'.(the account, session and run levels and the dry run answer the same sentence); - an unknown posture or mode;
shared_workspace: "shared_workspace no longer exists: list it as a workspace (folders: [{path, mode: "rw"}]). Nothing was saved.";- any other field: the earlier
allowed_folders,never_allowed,allow_any_folder,launch_folder_trust,mode,client_workspace_scope_overrides, … are all refused by name.
Rows may nest in any combination: the most specific row wins (the longest
real-path prefix, refused rows included). A refused /Users/me with an
allowed /Users/me/projects (rw) is valid, and so is a refused row inside an
allowed one.
The PUT answers like the GET. The GET also carries command_sandbox (round
12), the host's command sandbox state: {state:
"sandboxed"|"partial"|"unsandboxed"|"refused", kind, line, sentence,
unsandboxed_commands_allowed, configured, flag}. line is shown verbatim
(for example "Commands sandboxed: macOS sandbox-exec") and sentence is its
tooltip; see security.md.
Account: the account's default subset¶
GET /api/gateway/workspace/policy/{account} ({account} = name,
tenant:name, an entity's slug, or me) → {ok, policy: {account,
configured, posture, default_mode, folders}, gateway: <gateway policy>,
effective: <effective at the account level>, can_edit}. configured: false
means "follow the gateway policy"; the policy block then carries the gateway's
posture and default with no rows. 404 for an account that does not exist.
PUT /api/gateway/workspace/policy/{account}: {configured: false} follows
the gateway policy again; any of {posture, default_mode, folders}
configures the account (unnamed fields keep the stored value, else start from
the gateway's posture and default with no rows). A read-only or read & write
row must lie inside the eligible set (under "Deny everything…", under a listed
row; under "Allow everything…", outside every refused row and built-in
refusal) at most at its cap; otherwise the PUT is refused with "/workspace/policy/self answers 410
naming /workspace/policy/me.
Session: one conversation's subset¶
GET /api/gateway/sessions/{session_id}/workspaces → {ok, policy:
{session_id, account, configured, posture, default_mode, folders}, gateway,
account_default: <effective at the account level>, effective: <effective at
the session level>}. It works before the session's first run (nothing stored
= configured: false, "Use my default", whose display base is the account
default). PUT takes the account level's body and checks: {configured:
false} resets; rows are checked against the GATEWAY's eligible set, not the
account default (a conversation may use a workspace the account default does
not list). The gateway stores the choice in the owner's plane
(<plane data dir>/session_workspaces.json), so every app opening the
conversation sees it and it survives restarts. Only the session's owner, or an
admin with ?account=<tenant:name>, reads or changes it. Audited
(workspace_policy_changed, scope session).
Effective set¶
GET /api/gateway/workspace/effective/{account}[?session=<id>] → {ok,
account, session_id, level: "run" | "session" | "account" | "gateway",
posture, default_mode, folders: [{path, mode, cap, source: "gateway" |
"account" | "session" | "run"}], summary, gateway_summary}. level is the
level that applies (session > account > gateway); each mode is the lower of
the gateway's cap and the level's rule; default_mode is always ro or rw
(under "Deny everything…" it applies to nothing). summary and
gateway_summary are the lines every surface shows verbatim:
Deny everything, allow listed workspaces · /Users/me/Pictures (rw) · /Users/me/Documents (ro)
Allow everything, refuse listed workspaces (rw) · /secrets (refused) · /archive (ro)
The default mode appears in parentheses only under "Allow everything…".
POST /api/gateway/workspace/effective/{account} is a dry run: body
{workspace: {posture, default_mode, folders} | null, session?: <id>} →
the same shape (level: "run" for a payload; null = what a run would get).
Nothing is stored; a payload a run start would refuse is refused the same way.
Run start¶
Every door (POST /runs/start, schedules, host.start_run for bridges and
entities, automation occurrences) resolves a one-off workspace > the
session's choice > the account default > the gateway policy, clamps it to the
eligible set and caps, and sets the run's workspace arguments
(workspace_access_mode, allowed, read-only, writable and ignored paths). The
one-off is input_data.workspace (POST /runs/start also takes a top-level
workspace, moved there). With a session_id it wins for that run and is
saved onto the session when the session has no choice yet. At the HTTP doors a
row outside the eligible set or above its cap is refused (400
workspace_refused); the in-process doors drop or lower it, never widen, and
record each such row on the run as _gateway_workspace.clamped: [{path, asked,
got, sentence}] (got null = dropped; the sentence is the HTTP door's).
GET /runs/{run_id}/workspace also answers workspace_level,
workspace_summary and workspace_clamped. A
legacy workspace_allowed_paths list may only narrow; a workspace_root (a
launch folder, for example) must be reachable; a client workspace_access_mode:
"all_except_ignored" is refused. The run records the level under its vars
_gateway_workspace.{level, summary}. An automation keeps its choice in
target.input_data.workspace; a definition that sends the older
workspace_allowed_paths list gets it converted ("Deny everything, allow
listed workspaces", each workspace at its gateway cap). An automation saved
with Use my default (no payload) stores workspace: {"configured": false}
and follows its owner's default: each occurrence resolves session > account >
gateway when it is admitted (its run start), so a wider default applies from
the next run, with no revision. Nothing of the default is stored in the
definition (only a fail-closed fallback: the built-in denies and
workspace_access_mode: "workspace_only"); derived keys a client echoes back
with configured: false (an old snapshot) are ignored, never a narrowing. Every automation's workspaces, an explicit choice included, are guarded again at each occurrence against the policy as it is then: a workspace the admin refused or capped since is dropped or lowered for that run and recorded with its sentence (_gateway_workspace.clamped); the keys derived when the definition was saved are never used.
Artifacts and filesystem handoff¶
Gateway artifacts are the cross-package representation for files, media, and large payloads. Thin clients should pass artifact refs across runs instead of raw bytes or local paths:
{
"$artifact": "abc123",
"artifact_id": "abc123",
"run_id": "session_memory_sess-1",
"content_type": "image/png",
"filename": "input.png"
}
Gateway uses three distinct file-like source terms:
Artifact: a durable runtime-owned payload reference.Local File: a browser/client upload source. Hosted clients should upload bytes; browser-local paths are never interpreted as server paths.Server File/Server Folder: user-facing wording for a workspace-scoped server path under Gateway policy. The engineering contract is the canonicalWorkspacePathstring returned by/files/*, artifact import/export, and Runtime file nodes.
Hosted local uploads stay artifact-backed:
- one local file upload creates one artifact ref;
- multiple local files create an ordered list of artifact refs in Flow;
- a local folder uploads one artifact per file and may send
source_path(for examplereports/2026/summary.md) so relative member paths survive in artifact provenance without exposing browser-local absolute paths.
Upload a local file or folder member:
curl -sS -H "$AUTH" \
-F "session_id=sess-1" \
-F "source_path=reports/summary.md" \
-F "file=@./summary.md" \
"$BASE_URL/api/gateway/attachments/upload"
List run artifacts:
curl -sS -H "$AUTH" "$BASE_URL/api/gateway/runs/<run_id>/artifacts"
List artifacts visible to a session:
curl -sS -H "$AUTH" "$BASE_URL/api/gateway/sessions/sess-1/artifacts"
Browse server workspace files/folders:
curl -sS -H "$AUTH" \
"$BASE_URL/api/gateway/files/list?path=&include_directories=true&limit=200"
Optional filters:
- path: browse a specific workspace folder or mount alias.
- recursive=true
- family=image|video|audio|document|text|code|json|archive|other
- extensions=png,jpg or newline-separated values
- query=substring
- max_depth=<n>
Search artifacts across Gateway storage:
curl -sS -H "$AUTH" \
"$BASE_URL/api/gateway/artifacts/search?scope=all&artifact_kind=image&query=logo&tags=pin_id=image&include_stats=true&limit=500"
scope can be all, session, or run. Use session_id with
scope=session and run_id with scope=run; omit both for scope=all.
Search responses carry the row fields and also include
artifact_envelope_v1, a normalized projection of Runtime-owned descriptors,
access stats, and Gateway action links.
Useful query parameters:
- artifact_kind: UI-oriented kind filter. Comma-separated values match
semantic_kind, render_kind, or modality; generic audio means
unclassified audio and does not match canonical voice, music, or sound.
Single canonical kinds such as music, voice, image, markdown, or
json map to Runtime catalog filters. Multi-kind unions are supported, but
may be Gateway post-filters until Runtime exposes OR filters.
- semantic_kind / render_kind: canonical descriptor filters when the caller
wants the two dimensions separately.
- modality, content_type, workflow_id, node_id, created_after,
created_before, and tags: server filters for indexed descriptor fields.
- query: case-insensitive metadata search. Gateway may post-filter this field
when Runtime cannot index it directly.
- include_stats=true: include exact stats.total, byte totals, and facet
counts for the selected server-side filter set, independent of limit.
- limit, offset, and cursor: bounded paging. The default Runtime Explorer
page size is 500; limit<=0 is bounded unless debug_unlimited=true is used
by an admin/debug caller.
artifact_envelope_v1 contains normalized fields such as semantic_kind,
render_kind, workflow_id, node_id, turn_id, ledger_cursor,
generation, producer, media, source_refs, access, and links.
Sparse producer metadata is represented as missing fields; Gateway does not
invent provider/model provenance from filenames.
Generated-media artifacts created by child runs and projected into the parent run preserve Runtime descriptors and structured metadata. Direct transcription routes store transcript artifacts with source-audio refs, language/prompt hints, provider/model when available, and bounded route parameters. Descriptor-provided action links are sanitized to relative Gateway/UI links before they appear in envelopes; raw external provider URLs should be represented as trace availability or Gateway-owned trace records.
Content reads can label the access type for Runtime access stats:
curl -sS -H "$AUTH" \
"$BASE_URL/api/gateway/runs/<run_id>/artifacts/<artifact_id>/content?access_action=preview"
Supported access actions are content, preview, and download. The shorter
access=preview alias is also accepted.
Import a server workspace path into a session artifact:
curl -sS -H "$AUTH" -H "Content-Type: application/json" \
-d '{"session_id":"sess-1","source":{"kind":"workspace_path","path":"inputs/photo.png"},"pin_id":"image"}' \
"$BASE_URL/api/gateway/artifacts/import"
Export an artifact back into the server workspace:
curl -sS -H "$AUTH" -H "Content-Type: application/json" \
-d '{"path":"outputs/photo.png","create_parent_dirs":true,"overwrite":false}' \
"$BASE_URL/api/gateway/runs/<run_id>/artifacts/<artifact_id>/export"
Import and export use the same Gateway workspace policy as file helpers:
workspace roots, mounted roots, ignored paths, and size limits are enforced on
the server. Browser-local files should be uploaded through
POST /api/gateway/attachments/upload; browser-local file paths are not
interpreted as Gateway workspace paths. An upload belongs to the conversation
it was uploaded in: a run of another session that references it (any shape —
context.attachments, context.media, attachments, media, a bare id or one naming its
owner run_id) is refused at the start door with 400 artifact_not_in_session (and by the
host for in-process callers), so one conversation's attachment can never reach another's
model, whatever a client sends. Artifacts a run produced keep the hand-off below: a later
run, in any session of the same user, may name them by run_id. An artifact tagged
shared: user is visible to every session of its owner. In hosted user-auth mode, server
workspace import/export and /files/* helpers require an admin principal. They
never list, read or write the gateway data folder or the account's credential
folders (.ssh, .aws, Library/Keychains, ...), even when the workspace root
contains them (403).
Ordinary users can still upload browser-local files and list/search artifacts in
their own routed runtime.
Canonical Gateway server paths use rel/path for the main workspace root and
mount_alias/rel/path for approved mounts. When two allowed mounts share the
same basename, Gateway emits deterministic digest-suffixed aliases so the same
public path string can round-trip through /files/*, artifact import/export,
and Runtime file nodes.
Durable commands (POST /api/gateway/commands)¶
Commands are appended to a durable inbox and applied asynchronously by the runner.
Request fields (see SubmitCommandRequest in src/abstractgateway/routes/gateway.py):
- command_id: client-supplied idempotency key (UUID recommended)
- run_id: target run id (or session id for some event use-cases)
- type: pause|resume|cancel|conclude|emit_event|update_schedule|compact_memory|inject_guidance, or an automation.* type with run_id = the automation id (automations.md). Run commands aimed at an automation id answer 409 invalid_state.
- payload: command-specific object
Pause / cancel¶
curl -sS -H "$AUTH" -H "Content-Type: application/json" \
-d '{"command_id":"'"$(python -c 'import uuid; print(uuid.uuid4())')"'", "run_id":"<run_id>", "type":"pause", "payload":{"reason":"operator_pause"}}' \
"$BASE_URL/api/gateway/commands"
Resume a paused run¶
curl -sS -H "$AUTH" -H "Content-Type: application/json" \
-d '{"command_id":"'"$(python -c 'import uuid; print(uuid.uuid4())')"'", "run_id":"<run_id>", "type":"resume", "payload":{}}' \
"$BASE_URL/api/gateway/commands"
Resume a WAITING run with a payload (WAIT resume)¶
When payload.payload is present, the runner interprets this as “resume a WAITING run with a durable payload”:
curl -sS -H "$AUTH" -H "Content-Type: application/json" \
-d '{"command_id":"'"$(python -c 'import uuid; print(uuid.uuid4())')"'", "run_id":"<run_id>", "type":"resume", "payload":{"wait_key":"<optional_wait_key>", "payload":{"approved":true}}}' \
"$BASE_URL/api/gateway/commands"
Evidence: src/abstractgateway/runner.py (_apply_command, _apply_run_control).
Emit an external event¶
Minimal form:
curl -sS -H "$AUTH" -H "Content-Type: application/json" \
-d '{"command_id":"'"$(python -c 'import uuid; print(uuid.uuid4())')"'", "run_id":"<session_id>", "type":"emit_event", "payload":{"name":"chat.message","payload":{"text":"hi"}}}' \
"$BASE_URL/api/gateway/commands"
Evidence: src/abstractgateway/runner.py (_apply_emit_event).
Automations¶
An automation runs a workflow on a trigger (a fixed UTC interval, or only when asked) as a durable controller run whose occurrences are ordinary child runs. The full contract, with request and response shapes, the error envelope, typed waits, attention and operations, is in automations.md.
| Route | Answer |
|---|---|
GET /api/gateway/trigger-sources |
{items: [{id, version, label, capabilities, config_schema, event_schema, available, unavailable_reason?}]} |
POST /api/gateway/automations |
{request_id, title?, target, trigger?, context?, policy?} → {automation_id, revision, summary} |
GET /api/gateway/automations?status=&archived_only=&cursor=&limit= |
{items: [AutomationSummary], next_cursor, archived_automations} (older scheduled runs on the last page, legacy: true). Archived automations are left out unless archived_only=true (only them) or status names archived; archived_automations counts yours |
GET /api/gateway/automations/{automation_id} |
{definition, active_revision, summary} |
PATCH /api/gateway/automations/{automation_id} |
{command_id, expected_revision?, changes} → {command_id, accepted, duplicate, seq} |
POST /api/gateway/automations/{automation_id}/commands |
{command_id, type: automation.pause|resume|run_now|stop_current|archive|unarchive|revise, payload?} → the same receipt. unarchive brings an archived automation back paused with its history (409 invalid_state when it is not archived); an archived automation accepts only unarchive |
GET /api/gateway/automations/{automation_id}/occurrences?cursor=&limit= |
{items: [occurrence], next_cursor}, newest first |
GET /api/gateway/automations/{automation_id}/attention?cursor=&limit= |
{items: [attention item], next_cursor}: your unseen items, oldest first |
POST /api/gateway/automations/{automation_id}/seen |
{attention_cursor} → {attention_cursor} (moves forward only) |
POST /api/gateway/automations/{automation_id}/discuss |
{request_id, occurrence_index, prompt} → {session_id, run_id, session_kind: "discussion"} |
Errors on these routes, 401 and 403 included, always have the shape
{"detail": {"reason_code", "message", "field"?, "command_id"?}}
(automations.md). An occurrence's wait is answered
with the resume command above; the answer's shape depends on the wait's
kind (automations.md).
GET /api/gateway/runs rows carry session_kind, automation_id, role,
occurrence_index and legacy, plus workspace_root on turn rows (the folder
the run works in; absent on sub-runs), accept session_kind=chat,discussion, and
root_only=true returns conversation turns, one per occurrence
(automations.md).
With include_metrics=true each row also carries steps, llm_calls,
tool_calls and tokens_total: the totals of that run and every sub-run
below it, read from the ledger on either store backend (null when the gateway
has no ledger store). AbstractCode's conversation card shows the sum of its
turns' tool_calls.
Archived sessions¶
A conversation can be archived to take it out of the list. Archiving deletes
nothing: the session's runs, ledgers and artifacts stay, and
GET /api/gateway/runs?session_id=<id> still reads it (its rows then carry
archived: true and archived_at).
| Route | Answer |
|---|---|
POST /api/gateway/sessions/{session_id}/archive |
{session_id, archived: true, archived_at, archived_by, changed}; a repeat answers changed: false |
POST /api/gateway/sessions/{session_id}/unarchive |
{session_id, archived: false, archived_at: null, archived_by: null, changed} |
GET /api/gateway/runs?root_only=true |
leaves archived sessions out; every listing carries archived_sessions (how many sessions of the listed kind you have archived) |
GET /api/gateway/runs?root_only=true&archived_only=true |
only the turns of archived sessions, each with archived: true |
Session purpose (kind). A session is a conversation (the default) or a
docs chat (a Docs assistant conversation). POST /api/gateway/runs/start
takes kind: "docs" with a session_id (the kit's Docs assistant sends it);
any other value, or docs without a session, is a 400. Turn listings
(GET /runs?root_only=true or archived_only=true, without session_id /
parent_run_id) take kind=conversation (the default), kind=docs or
kind=all, so a docs chat never appears in a conversation list while staying
in the same pool: readable by session_id, ledgered, archivable. Every row
carries kind. The mark lives in the plane's session_kinds.json; a session
without one is a conversation. A run of the shipped docs-qa workflow marks
its session docs even without kind, and the first turn listing runs a
one-time migration (docs_qa_bundle_v1) that marks every existing session
whose turn root ran docs-qa (matched on the workflow's bundle id).
The session's owner or an admin may archive it. Each account works in its own
runtime plane, so a session that belongs to another account answers
404 {"detail": {"reason_code": "session_not_found"}}. The mark and its
history (every archive and unarchive) are kept in the plane's
session_archive.json; each call writes a session.archived /
session.unarchived audit event. Automations are archived with the
automation.archive command and brought back with automation.unarchive.
Beyond the core¶
/api/gateway/* also includes optional operator/tooling endpoints (reports inbox, triage queue, backlog browsing + exec runner, process manager, file/attachment helpers, embeddings, voice, discovery, …).
See: maintenance.md.
Discovery endpoints (optional)¶
These exist to help thin clients adapt to the deployed gateway.
- Capabilities (best-effort):
GET /api/gateway/discovery/capabilities - Providers/models discovery (best-effort):
GET /api/gateway/discovery/providers,GET /api/gateway/discovery/providers/{provider}/models - Tools (thin-client allowlist help):
GET /api/gateway/discovery/tools. Round 12: each process-spawning tool row (AbstractRuntime'sSANDBOXED_TOOL_NAMES:execute_command,shell_exec,local_helper_start,execute_python) carriessandboxed: true|falseandsandbox(the state sentence, for example "Sandboxed to this run's workspaces"), and the answer carriescommand_sandbox(as onGET /workspace/policy) - Skills inventory:
GET /api/gateway/skills— the abstractskill shelf with trust verdicts (roster rows{name, description, trust_level, blocked, requires_review, tree_hash, source, has_scripts, reasons}); degradations are labeledwarnings, never a fabricated list. The shelf is, in order: the saved settingskills.shelf, then the legacy launch environment value (reported as such), then the gateway's own copy in<data dir>/skills/registry, then a framework checkout's shelf (only when that copy is missing), else none; the response names which asshelf_source(stored,env,seeded,checkout,none). See Skills shelf. - MCP server inventory:
GET /api/gateway/mcp/servers— the declared registry at<data_dir>/config/mcp_servers.json({"version": 1, "servers": [{"name", "url"?, "description"?, "auth_required"?, "tags"?}]}), served with declared fields only andprobed: false(connect state/tool counts require a probe lane and are never faked). - Dynamic capability catalogs:
GET /api/gateway/voice/voices,GET /api/gateway/audio/speech/models,GET /api/gateway/audio/transcriptions/models,GET /api/gateway/audio/music/providers,GET /api/gateway/audio/music/models,GET /api/gateway/vision/provider_models
The capabilities payload includes package presence (abstractruntime,
abstractcore, abstractmemory, abstractvoice, abstractvision), existing
gateway helpers (tools, visualflow, media), memory-store readiness, and
AbstractCore capability plugin status for voice, audio, vision, and
music.
The route paths and contract descriptors are the stable part of this surface. Catalog routes also include a stable Gateway-owned envelope:
catalog.contract = gateway_catalog_v1catalog.version = 1items = [...]
The lower-layer fields stay in the payload for compatibility. Thin clients
should read catalog plus items; the route-specific fields (models,
provider_models, profiles, voices) remain available.
Provider discovery also reports the resolved default provider/model when one is
configured. The resolver follows request values, flow pins, and the execution-host
input.text capability route; if no pair exists, the response includes
default_error rather than a hardcoded local model.
It also includes a versioned thin-client contract:
capabilities.contracts.version: currently1capabilities.contracts.common: shared run start/list/summary/input/history, ledger, artifact, attachment, workspace, discovery, provider prompt-cache controls, and the host-visibility descriptorsmodel_residency(includingrow_schema = "model_residency_row_v1"and the canonicalmodality_uicolor map),host_state, andsession_caches(see Host state and model residency).common.artifactsincludes run listing/content, session artifact listing, artifact search withartifact_envelope_v1, exact stats/facets,artifact_kindUI filtering, workspace import, and workspace export descriptors when available. Permission-sensitive descriptors are principal-aware: ordinary users see admin-only workspace import/export and provider prompt-cache controls marked unavailable withadmin_requiredmetadata.capabilities.contracts.common.readiness: compact Gateway-ownedgateway_surface_readiness_v1summary derived from the shared endpoint/media/ residency descriptorscapabilities.contracts.flow_editor: the AbstractFlow editor/runtime surfacecapabilities.contracts.assistant: assistant-facing voice/audio/media/cache feature gatescapabilities.contracts.abstractcode: code-client run/history/workspace/cache feature gates
Contract booleans are intentionally conservative. Package installed=true is
not the same thing as endpoint available=true; clients should branch on the
versioned contract fields when enabling controls.
common.readiness is intentionally narrower than provider/backend health. It
summarizes Gateway surface availability from existing descriptors, but it does
not invent selected backend/provider/model truth or stable degraded-state
reason codes.
Evidence: src/abstractgateway/routes/gateway.py (discovery_capabilities, discovery_providers).
AbstractFlow gateway-first editor contract¶
The browser editor can use AbstractGateway as its runtime and storage host.
Draft VisualFlow records:
GET /api/gateway/visualflowsPOST /api/gateway/visualflowsGET /api/gateway/visualflows/{flow_id}PUT /api/gateway/visualflows/{flow_id}DELETE /api/gateway/visualflows/{flow_id}POST /api/gateway/visualflows/{flow_id}/publish
Bundle inspection and editor run-schema helpers:
GET /api/gateway/bundlesGET /api/gateway/bundles/{bundle_id}GET /api/gateway/bundles/{bundle_id}/flows/{flow_id}GET /api/gateway/bundles/{bundle_id}/flows/{flow_id}/input_schema
Workflow description (the console's inline edit on the Workflows page):
| Route | Body → answer |
|---|---|
PATCH /api/gateway/bundles/{bundle_id} |
{description} (at most 2000 characters; "" goes back to the file's own description) → {ok, bundle_id, owner, description, description_edited, updated_by, updated_at} |
Only the owner may: a user for their own workflows, an admin for the gateway's.
A user on a shared workflow gets 403 admin_required, another user's workflow
is 404, a shipped workflow 409 workflow_shipped. The .flow file is never
rewritten (Export keeps the original bytes); the text applies to every version
and is kept next to the owner's archive file (config/workflow_descriptions.json).
Each attempt is audited as workflow.description (lengths, never the text).
GET /api/gateway/bundles items carry the effective description,
description_edited and actions.can_edit_description; the envelope's
interfaces ({<interface id>: {label, help, known}}) names every interface
the listed workflows declare, readable by any signed-in account (the console's
"Used by" column).
The input-schema endpoint returns a versioned payload with:
versionbundle_id,bundle_version,bundle_ref,flow_id,workflow_idinputs: entrypoint input pins derived from theon_flow_startnodedefaults: pin defaults from VisualFlow JSONinput_data_schema: a small JSON Schema object for the Run Flow modal
Example:
curl -sS -H "$AUTH" \
"$BASE_URL/api/gateway/bundles/my-bundle/flows/ac-echo/input_schema"
Native-loop bundles (react / codeact / memact)¶
Some shipped bundles declare metadata.native_loop_factory instead of VisualFlow
JSON (manifest.flows is empty). The gateway materializes an abstractagent
loop at load time. Discovery uses the same bundle list endpoint — not
/discovery/workflows.
Thin clients should:
GET /api/gateway/bundles(authenticated).- Filter entrypoints whose
interfacesincludesabstractcode.agent.v1. - Read
metadata.native_loop_factory(react,codeact, ormemact) to distinguish native loops from VisualFlow agent bundles. - Use each entrypoint's
workflow_id(for examplereact-agent@0.1.0:react). - Start runs with
POST /api/gateway/runs/startandbundle_id/flow_idfrom the bundle listing (for examplereact-agent+react).
Native-loop entrypoints do not ship VisualFlow JSON. The gateway serves a
versioned input-schema stub (prompt required; provider and model
optional) from
GET /api/gateway/bundles/{bundle_id}/flows/{flow_id}/input_schema.
Headless clients may also pass those fields without fetching the schema.
The shipped react-agent@0.1.0 bundle is built by
scripts/build_react_agent_bundle.py and force-included in the wheel. A running
gateway process must restart (or call bundle reload) after the file lands on
disk before /bundles lists it.
Run history bundle (GET /runs/{run_id}/history_bundle)¶
Thin clients should prefer this endpoint over stitching ledger, session, and artifact endpoints. The export is owned by AbstractRuntime; the gateway forwards query parameters and returns the bundle JSON unchanged (including in-band degradations).
Query parameters:
| Parameter | Default | Notes |
|---|---|---|
include_subruns |
true |
Descendant runs in the bundle tree |
include_session |
false |
Root session turn list |
session_turn_limit |
200 |
Cap when include_session=true |
ledger_mode |
tail |
tail or full |
ledger_max_items |
2000 |
Per-run ledger cap when ledger_mode=tail |
detail |
full |
full (complete payloads) or replay (transcript-fold projection) |
detail=replay drops request-side payloads and observability paths the
transcript fold never reads; runtime marks each omission with $omitted inside
ledger records. Use it for session replay and thin-client folds — it is much
smaller than full (gzip helps further; send Accept-Encoding: gzip).
warnings (always present, may be empty): typed degradations the export
survived instead of failing silently. Each entry is an object with at least
code and detail; many include run_id. Known codes today:
| Code | Meaning |
|---|---|
subtree_discovery_failed |
Could not list child runs; bundle covers root only |
subtree_truncated |
Descendant discovery hit the run cap |
ledger_read_failed |
Ledger for a run id could not be read |
torn_rows_skipped |
Corrupt/unparseable ledger lines skipped |
ledger_tail_window |
Ledger truncated to ledger_max_items (tail mode) |
input_data_offload_failed |
Input-data artifact reference could not be resolved |
Clients must surface non-empty warnings to the operator — a bundle that
"looks complete" but carries warnings may be missing subruns, ledger tail, or
offloaded input data.
Session history bloc (GET /sessions/{session_id}/history/bloc)¶
Returns one cursor-bounded bloc of root session turns, each with an inline
history_bundle export — one round-trip instead of N per-turn bundle fetches. Resume pagination uses an ISO created_at cursor in the
before query parameter (never turn-count offsets).
The turns are the ones GET /api/gateway/runs?root_only=true lists for the
session: parent-less runs plus automation occurrences (a retried occurrence
once, as its last attempt), never automation controllers
(automations.md).
Query parameters:
| Parameter | Default | Notes |
|---|---|---|
before |
(omit) | ISO-8601 cursor; only turns strictly before this timestamp |
limit |
5 |
Max turns in this bloc (1–50) |
detail |
replay |
Forwarded to each turn's bundle export (full | replay) |
include_subruns |
true |
Per-turn bundle tree |
ledger_mode |
tail |
tail or full |
ledger_max_items |
2000 |
Per-turn ledger cap when ledger_mode=tail |
include_drafts |
false |
Include draft-test root runs |
Response fields: session_id, cursor_before (echo of before), cursor_after
(oldest turn returned — pass as the next before), older_remaining, warnings,
and turns[] (run_id, created_at, status, bundle or error).
The editor observes runs with the core lifecycle endpoints above:
/runs/start, /runs/{run_id}, /runs/{run_id}/ledger,
/runs/{run_id}/ledger/stream, /runs/ledger/batch,
/runs/{run_id}/input_data, /runs/{run_id}/history_bundle, and
/runs/{run_id}/artifacts.
Optional multimodal scope¶
Current direct Gateway endpoints:
- POST /api/gateway/runs/{run_id}/voice/tts
- POST /api/gateway/runs/{run_id}/voice/tts/stream
- POST /api/gateway/runs/{run_id}/audio/transcribe
- POST /api/gateway/runs/{run_id}/images/generate
- POST /api/gateway/runs/{run_id}/images/edit
- POST /api/gateway/runs/{run_id}/images/upscale
- POST /api/gateway/runs/{run_id}/videos/generate
- POST /api/gateway/runs/{run_id}/videos/from_image
- POST /api/gateway/runs/{run_id}/music/generate
- GET /api/gateway/voice/defaults
- GET /api/gateway/voice/voices
- GET /api/gateway/audio/speech/models
- GET /api/gateway/audio/transcriptions/models
- GET /api/gateway/audio/music/providers
- GET /api/gateway/audio/music/models
- GET /api/gateway/vision/provider_models
- GET /api/gateway/vision/adapters
/voice/defaults is the one answer to "which engines speak and listen by default":
{"tts": {"route": "output.voice", "configured": true, "provider": "supertonic", "model": "supertonic-3", "voice": "M3"},
"stt": {"route": "input.voice", "configured": true, "provider": "faster-whisper", "model": "large-v3"},
"source": "capability_defaults"}
A route the administrator has not set reads {"configured": false, "provider": null, "model": null, "note": "No gateway default is set for …"}. A synthesis or transcription request that names no provider and no model runs exactly these routes, so apps show "Gateway default · supertonic / supertonic-3" from here — never from the voice catalog's engine-side fields. The catalog repeats the answer (gateway_defaults; active_tts_provider / active_stt_provider = the configured routes, absent when unset). /audio/transcribe returns the route that ran (provider, model) and duration_ms; a language hint skips the engine's language detection (faster-whisper large-v3 on an Apple-Silicon CPU: ~25 s → ~9 s for a 4 s clip).
/voice/tts returns a durable audio artifact after synthesis. /voice/tts/stream
returns JSON Lines stream events for progressive playback when discovery advertises
capabilities.contracts.assistant.voice.tts.streaming=true; successful streams still
finish with a Runtime-owned child-run audio artifact. The stream is real streaming: the engine splits the text at sentence boundaries (first segment one short clause), each segment is sent as soon as it is synthesised while the next one is synthesised, and the setup runs off the event loop. The terminal done event's metrics carry ttfb_s, rtf, device and, on the CPU, device_reason.
Stream lines (JSON Lines): runtime_start (run_id = the run named in the path, child_run_id = the id the outcome will be recorded under, schema: abstractruntime.tts_stream.start.v2), engine events (start, audio with sequence and audio_b64), then one terminal line: done (with audio_artifact and child_run_status), cancelled or error (error = the sentence; watchdog_timeout: true when the engine went silent past ABSTRACTGATEWAY_VOICE_TTS_TIMEOUT_S). When the engine is busy (another stream, or the model using the machine) and its setup takes more than a second, the first line is {"type": "queued", "message": "Waiting for the voice engine: …"}; speech follows when the engine is free. A request past the voice synthesis bound (ABSTRACTGATEWAY_VOICE_MAX_CONCURRENCY) waits one second, then gets 503 "Read aloud is busy: N voice syntheses are already running on this gateway …".
Reading aloud never changes the run's state: no wait is created while audio streams, and the child run is recorded already completed (done, or with errors[0].code cancelled/stream_error) when the stream ends. A stream cut by a gateway restart leaves no run behind; a wait left by a gateway up to 0.13.0 is closed at startup with errors[0] = {"code": "interrupted", "message": "Read aloud was interrupted because the gateway restarted; …"}.
The catalog endpoints proxy AbstractCore Server routes when
ABSTRACTCORE_SERVER_BASE_URL
is configured. Gateway uses explicit Core auth settings for that hop and never
reuses the Gateway bearer token as a Core/provider secret. Without a configured
Core server, the voice/model routes return bounded static descriptors from
Gateway and capability-package environment variables.
Each route adds:
catalog: Gateway-owned route metadata (contract,version,kind,scope,route_source, optionalupstream_source, and route filters)items: one canonical primary array for thin clients
Examples:
- Voice provider listings (
/voice/voices?providers_only=true,/audio/speech/models?providers_only=true,/audio/transcriptions/models?providers_only=true) always list the cloud providersopenaiandopenai-compatible; theiritemscarryneeds_key,key_source(environment|providers| null),stateandreason, andcloud_providersrepeats them. A key counts when it is in the environment or saved through the Providers screen. Local engines are listed when AbstractVoice reports their runtime installed;unavailable_providers(per kind:{provider, code: runtime_missing|model_not_downloaded|not_configured, reason}) andunavailable_reasonsay why others are missing, and a listing filtered to a cloud provider without a key says where the key goes. /voice/voices:itemscontain voice/profile records withid,label, optionalprovider, optionalmodel, andvoice_kind/audio/*/models:itemscontain model records withid,label, optionalprovider, optionaltasks, and optionalparameters/audio/music/providersand/discovery/providers:itemscontain provider records withid,label, andprovider
Generated images are available through Runtime workflows when a compatible
image backend is installed and configured. Gateway also exposes a direct image
generation endpoint that uses the Runtime/Core output-selector contract rather
than a provider-specific image client. The route creates a durable child run,
stores the generated image as a run artifact, and returns
event_name="abstract.progress" so thin clients can stream the child-run ledger
for progress:
run_id,request_id,prompt- optional
provider,model,size,width,height,format, batchcount/n,seeds, and orderedlora_adapters image_artifact: first generated image for compatibilityimage_artifacts: full ordered image artifact list for batch generation
size, width, and height are optional passthrough request overrides. Do
not inject a client-side default size. Different image providers/models accept
different size sets; when the client leaves dimensions unset, Runtime/Core lets
the configured backend use its default or auto behavior.
If the active workflow runtime already has an AbstractCore LLM client, the route
uses it. For tools-only workflows, the route can create a direct Runtime/Core
client from request provider/model or the execution-host capability route
default. Unsupported or unconfigured deployments return a structured ok=false
response instead of a failed run.
Gateway also exposes a direct image-edit sibling route:
POST /api/gateway/runs/{run_id}/images/edit
The request uses a source image_artifact, optional mask_artifact, the same
provider/model and image backend selectors as image generation, plus optional
batch count / n, seeds, and ordered lora_adapters, and returns an
artifact-backed edited image. Batch responses also return image_artifacts.
Thin clients should feature-detect it from
capabilities.contracts.flow_editor.media.edited_image or
capabilities.contracts.assistant.media.edited_image. It uses the same
child-run abstract.progress progress contract as direct image generation.
Gateway also exposes a direct image-upscale sibling route:
POST /api/gateway/runs/{run_id}/images/upscale
The request uses a run-visible source image_artifact, optional provider/model
selectors, and optional upscaler controls such as scale, resolution,
softness, seed, quantize, and vae_tiling; resolution may be a
shortest-edge integer or a scale factor such as 2x. Thin clients should
feature-detect it from capabilities.contracts.flow_editor.media.upscaled_image
or capabilities.contracts.assistant.media.upscaled_image, list models with
GET /api/gateway/vision/provider_models?task=image_upscale, and stream the
returned child-run ledger for abstract.progress events.
Generated music follows the same direct child-run pattern. Thin clients should
discover it from capabilities.contracts.flow_editor.media.generated_music or
capabilities.contracts.assistant.media.generated_music, list providers/models
from the music catalog routes, and treat the returned child_run_id plus
music_artifact as the durable output handle.
Generated video also follows the direct child-run pattern:
POST /api/gateway/runs/{run_id}/videos/generateuses the Runtime/Coreoutput.modality=video/task=text_to_videocontract and accepts optional batchcount/n,seeds, orderedlora_adapters, andflow_shift.POST /api/gateway/runs/{run_id}/videos/from_imageaccepts a run-visible sourceimage_artifact, accepts the same optional batch/adapter/video control fields, and usestask=image_to_video.- Thin clients should discover these routes from
capabilities.contracts.flow_editor.media.generated_videoandcapabilities.contracts.flow_editor.media.image_to_video(or the matchingassistant.media.*entries), useGET /api/gateway/vision/provider_models?task=text_to_video|image_to_videofor model catalogs, useGET /api/gateway/vision/adaptersfor compatible installed adapter catalogs, stream the returnedchild_run_idledger forabstract.progressevents, and readvideo_artifactswhen batch generation is requested.
STT and listen contract notes:
POST /api/gateway/runs/{run_id}/audio/transcribeaccepts a run-visibleaudio_artifactplus optionallanguage,prompt,response_format,temperature,format,provider, andmodelhints.capabilities.contracts.flow_editor.voice.sttandcapabilities.contracts.assistant.voice.sttpoint to that upload route.capabilities.contracts.flow_editor.voice.listenandcapabilities.contracts.assistant.voice.listenare host-capture contracts, not a live microphone socket. They tell higher apps to capture locally and emit an event or upload the resulting audio artifact.
KG memory¶
POST /api/gateway/kg/query queries the configured AbstractMemory TripleStore.
Gateway resolves the store through:
ABSTRACTGATEWAY_MEMORY_STORE_BACKEND=lancedb|memory(sqlitewhen the installed AbstractMemory build exposesSQLiteTripleStore)ABSTRACTGATEWAY_MEMORY_STORE_PATHABSTRACTGATEWAY_MEMORY_REQUIRE_VECTOR
Structured queries work with LanceDB and in-memory stores. SQLite also works
when the installed AbstractMemory build exposes SQLiteTripleStore. Semantic
query_text requires a vector-capable backend plus the execution-host
embedding.text route; SQLite returns a clear 400 instead of pretending to
support semantic recall.
Capability discovery reports KG memory as available when AbstractMemory is installed and the configured backend can be resolved. A fresh persistent store does not need to exist yet; empty-store structured queries return an empty result rather than making Flow authoring nodes unavailable.
Models and engines¶
The gateway serves AbstractCore's models and engines payloads unchanged,
under /api/gateway. The bodies and payloads
are the same as AbstractCore's own /acore/* routes; abstractcore and
abstractgateway render them with the same screens.
| Method and path | Access | Body / query | Returns |
|---|---|---|---|
GET /host/profile |
user | refresh=1 |
host_profile_v1 |
GET /engines |
user | probe=1 |
gateway_engines_v2 rows (AbstractCore's detection plus the install plan and actions) with install_allowed and install_policy; see engines.md |
GET /engines/{id} |
user | probe=1 |
one engine row plus install_allowed; 404 for an unknown id |
POST /engines/{id}/install |
admin | {"dry_run": bool, "force": bool, "location": "auto"\|"user"\|"system"} |
an engine_install_job_v1 job (user-level first; pauses in needs_admin / needs_tools), see engines.md |
GET /engines/jobs, GET /engines/jobs/{id} |
user | engine install jobs | |
POST /engines/jobs/{id}/continue, /cancel |
admin | {"action"?} |
the job |
POST /engines/{id}/start, /stop |
admin | Ollama / LM Studio server state | |
GET /models/catalog |
user | q, engine, fits=1, hub=1, tag (repeatable) |
model_catalog_v1 |
GET /models/installed |
user | provider |
models_installed_v1 |
POST /models/download |
admin | {"provider", "artifact", "dry_run", "expected_bytes"?} or {"recommended": true} |
{"ok": true, "job": {...}}; with recommended, {"ok": true, "recommended": true, "jobs": [...], "group": {...}} |
GET /models/download/{job} |
user | {"ok": true, "job": {...}} (a grp_... id returns the parent job) |
|
GET /models/downloads |
user | {"ok": true, "jobs": [...]}, newest first, parents included |
|
POST /models/download/{job}/cancel |
admin | none | {"ok": true, "job": {...}}; stops the transfer within about a second; a grp_... id cancels every running child; 404 when unknown |
GET /models/downloads/stream |
user | job_id, until_idle=1 |
Server-Sent Events of the same dicts, see model-downloads.md |
POST /models/delete |
admin | {"provider", "artifact", "dry_run": bool, "force": bool} |
host_job_v1 (kind delete) |
POST /models/delete-download |
admin | {"provider", "artifact", "dry_run": bool} (the catalog's own names) |
model_download_delete_v1, see "Delete a download" below |
GET /jobs |
user | kind, status |
{"schema": "host_jobs_v1", "jobs": [...], "generated_at"}, newest first |
GET /jobs/{id} |
user | host_job_v1; 404 when unknown |
|
POST /jobs/{id}/cancel |
admin | none (an empty {} is accepted) |
host_job_v1; 404 when unknown |
Empty query values (q=, engine=) mean "no filter". probe, fits and
hub accept 1/0 and true/false.
Catalog artifacts. model_catalog_v1 is AbstractCore's payload, served
unchanged (field reference: AbstractCore docs/models.md, "The catalog").
Besides quant (the artifact's own label, lowercased, or null) and bits
(effective bits per weight), every artifact carries quant_class, one of
2bit, 3bit, 4bit, 5bit, 6bit, 8bit, 16bit, full, unknown,
for filtering by quantization: q4_k_m, 4bit, mxfp4 and oq4e are
4bit; q8_0 and 8bit are 8bit; bf16 and f16 are 16bit; f32 is
full. quant_class_source is stated (the reference names its quant),
assumed (a bare Ollama tag such as qwen3.5:9b or LM Studio id: the class of
the engine's default build, which the fit estimate assumes too) or null (no
quant information; the class is unknown). options holds the route options a
recommendation copies with the artifact ({} for most); companions lists
repos downloaded with it (an MLX build's MTP drafter, from AbstractCore's
drafter registry; [] for most), companion_bytes is their size, and
download_bytes already includes it; note is one sentence about the build. On Apple silicon the text rows pre-select the
memory tier's MLX build and exactly one text row is the starter.
Jobs. A host_job_v1 has schema, job_id, kind
(download | delete | engine_install), status
(queued | running | completed | failed | cancelled), provider, artifact,
engine, percent, downloaded_bytes, total_bytes, message, log_tail,
command (the exact argv), dry_run, started_at, finished_at, error
(a string or null), joined, result and cli_equivalent, which names the
abstractgateway command that does the same thing. A dry run finishes before
the POST returns. On the /models/download routes the job also carries
job (the id), events, host_status, reports queued as running, and
counts joined including the first request.
Download progress. A download job also carries state
(queued | resolving | downloading | verifying | installing | done | failed |
cancelled | stalled), bytes_done, bytes_total, size_unknown,
size_note, bytes_per_second, eta_s, updated_at, files
([{name, bytes_done, bytes_total, state}]), current_file, a one-sentence
message, the tool's own detail, and transitions. "Use recommended
defaults" ({"recommended": true}) returns one parent job (kind:
"download_group", id grp_...) whose bytes, percent, speed and time left
add up its children. The full contract, one real example per state and what
each source reports: model-downloads.md.
Refusals share one body:
{"ok": false, "status", "reason"?, "message", "detail", "error": {"message", "type"}, ...}.
| Status | When |
|---|---|
400 invalid |
provider or artifact missing |
403 refused / not_allowed |
a real engine install while allow_engine_install is off (configuration.md); install_policy says why |
| 403 | the caller is not an admin (every POST above) |
404 not_found |
unknown job id, engine id, or a model that is not installed |
409 busy |
an engine install is already running (job is the running one) |
409 refused |
the engine is not supported here or has no install command (install is the plan), or a delete is blocked (delete_blockers: loaded, shared_cache:…, unknown_location, engine_not_running, remote_engine; force: true overrides the first two) |
501 unsupported / abstractcore_too_old |
the installed AbstractCore is too old for these routes; required, installed and missing name what to upgrade |
503 unavailable |
AbstractCore is not installed |
Delete a download (POST /models/delete-download, the console's
Models-page Delete). Synchronous; removes one downloaded artifact with its
engine's own mechanism: Ollama's DELETE /api/delete (what ollama rm does),
the Hugging Face / MLX cache folder of the repo, or, for org/repo:QUANT
(a llama.cpp GGUF quant), only that quant's files (the repo goes when it was
the last model file set). dry_run: true deletes nothing and answers the
exact bytes for the confirmation. Answers:
| Status | Body |
|---|---|
| 200 | {"schema": "model_download_delete_v1", "ok": true, "status": "planned" \| "deleted", "provider", "artifact", "freed_bytes", "paths", "command", "also_used_by", "presence": "installed" \| "absent", "message"} |
409 refused |
reason: resident (this gateway or the engine has it loaded: "Unload it first"), locked ("Unlock and unload it first"), downloading (its download is still running), managed_elsewhere (LM Studio: its CLI has no remove command, delete it in LM Studio), engine_not_running, unknown_location, remote_engine; message and fix are sentences to show as they are |
404 not_found |
reason: "not_downloaded" |
502 failed |
reason: "engine_failed": the engine did not delete; command names what ran |
A build whose files AbstractCore files under the sibling engine of the shared
Hugging Face cache (MLX vs Hugging Face) is deleted all the same, and
also_used_by names that engine. Every real delete writes
model.download_deleted (provider, artifact, actor, freed_bytes,
paths) to <data_dir>/audit_log.jsonl; every refusal (a dry run's too)
writes model.download_delete_refused with its reason and dry_run.
Example:
curl -s -H "Authorization: Bearer $TOKEN" "$GW/api/gateway/models/catalog?q=qwen3&fits=1" | jq '.rows[0].artifacts[0].fit'
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"dry_run": true}' "$GW/api/gateway/engines/ollama/install" | jq '.command, .cli_equivalent'
Host state and model residency¶
Gateway exposes a host-level view of the execution machine — memory, GPU, resident models, and session prompt caches — so consoles and agents can render an "agentic OS" panel from one API surface.
Read endpoints (any authenticated principal):
GET /api/gateway/host/state— one-call host snapshotGET /api/gateway/host/metrics/memory— host memory snapshotGET /api/gateway/host/metrics/gpu— GPU utilization probeGET /api/gateway/models/loaded— model residency listingGET /api/gateway/models/context_estimate— context/KV memory estimate for a provider+modelGET /api/gateway/sessions/prompt_cache— session prompt-cache enumeration
Mutation endpoints (admin principal required):
POST /api/gateway/models/load— load (and by default pin) a model runtimePOST /api/gateway/models/unload— unload a model runtimePOST /api/gateway/models/lock— lock a resident model against unloadPOST /api/gateway/models/unlock— release a model-residency lockPOST /api/gateway/models/download— fetch model weights onto the hostPOST /api/gateway/sessions/{session_id}/prompt_cache/clear_all— clear every runtime-minted prompt cache for a session
Reads are visibility every authenticated client needs; mutations spend shared
host resources and stay operator acts. Anonymous requests are rejected on all
of these routes, like every other /api/gateway/* path.
GET /host/state¶
One snapshot with memory, gpu, models, and session_caches sections:
{
"ok": true,
"ts": 1787857000.0,
"memory": {"ram": {"...": "..."}, "process": {"rss_bytes": 140443648}, "device": {"backend": "metal", "allocated_bytes": 0, "...": "..."}},
"gpu": {"supported": true, "source": "ioreg", "gpus": [{"name": "...", "utilization_gpu_pct": 0.0}]},
"models": [{"runtime_id": "...", "provider": "...", "model": "...", "resident": true, "...": "..."}],
"session_caches": [],
"totals": {"models": 2, "models_resident": 1, "model_bytes": 3109915433, "session_caches": 0, "session_cache_bytes": null},
"degraded": [],
"row_schema": "model_residency_row_v1"
}
modelsrows use the frozenmodel_residency_row_v1schema described below;session_cachesrelays the runtime facade's cache rows verbatim.- Every section is independently best-effort and the route never returns a
- A missing facade method or a failed probe nulls that section and names
it in
degraded; areasonsmap (present only when non-empty) says why. - The
gpusection keeps its in-band{"supported": false, "reason": "..."}payload when the probe answers but reports no support; it still counts as degraded. totals.model_bytessums the knownsize_bytesvalues and isnullwhen no row reports a size;totals.session_cache_bytesbehaves the same over the cache rows'bytes.memory.devicerelays AbstractCore's accelerator figures, includingprocess_held_bytes(what this gateway process holds across its model libraries: MLX, llama.cpp, transformers, embeddings) andprocess_held_basis, how that figure was measured (metal_device_counter,cuda_device_counter…, orsum:of the libraries' own counts).memory.resident/memory.heldname the models that hold it (backend, model, number of holders) when the runtime reports them.residency_diagnosticsis always present ({}when the runtime reports nothing): the dictGET /models/loadedreturns asdiagnostics, withpending_ejects(models that will be ejected when their in-flight call ends) andlast_switch_ejects(what the last default-model switch ejected, kept or failed to eject, with the reason). The consoles and the tray show these as "Will eject X when the in-flight call ends", "X: eject failed: reason".totals.modelscounts every known row — configured / cached rows included — whiletotals.models_resident(additive) counts only rows withresident: true. Clients that display "N loaded" must readmodels_resident: default ≠ loaded, and presenting configured capability defaults as loaded is exactly the lie this field removes.- When the runtime memory snapshot reports a host identity, the response also
carries a top-level
hostobject (the identity facts of the machine the snapshot describes). The block is omitted when the runtime does not report one. Together with the per-rowhost_id/host_namefields below, this is the seam a multi-machine resource pool would aggregate on; one gateway binds one runtime host, and the pool design is proposed in backlog 0093.
GET /host/metrics/memory¶
Returns {"ok": true, "supported": true, ...} plus the snapshot sections:
ram (total/available/used bytes and percent), process (rss_bytes), and
device (backend, allocated_bytes, total_bytes, free_bytes). When the
runtime host facade does not expose a memory snapshot, the route answers 200
with {"ok": true, "supported": false, "reason": "..."} — the same degraded
style as GET /host/metrics/gpu.
How to compare memory measurements: process.rss_bytes and
device.allocated_bytes are different axes. In-process device backends (for
example MLX on Metal) return freed buffers to the process heap and the
operating system may retain those pages, so process RSS does not shrink when a
model unloads. Use device.allocated_bytes to verify that an unload freed
device memory; use ram and process for overall host pressure.
Model residency (/models/loaded, /models/load, /models/unload)¶
GET /models/loaded lists the model runtimes the host knows about, with
optional task, provider, model, and base_url query filters. The
response keeps the raw runtime records in models and adds a normalized
rows array in the frozen model_residency_row_v1 schema (named by
row_schema), so thin clients do not need per-provider alias tables.
Each model_residency_row_v1 row has exactly these fields (unknown values
are null, never guessed):
runtime_id, task, provider, model, source, resident, state,
pinned, default, size_bytes, size_vram_bytes, expires_at,
context_length, loaded_at, last_used_at, locked, lockable,
modalities, calibrated_context_length, context_calibrated, host_id,
host_name, details
The schema is additive-tolerant and keeps the model_residency_row_v1 name
as optional fields are added; treat fields beyond the original 16 as
optional. The lock/calibration/host fields mean:
locked/lockable— tri-state booleans: whether the model is locked against unload, and whether this runtime supports locking it at all.modalities— list of modality strings when the runtime reports one (nullotherwise, including when the value is not a clean string list).calibrated_context_length/context_calibrated— the measured usable context length and whether it came from calibration rather than metadata.host_id/host_name— identity of the machine serving the model, stamped by the runtime that reported the row (see thehostblock note underGET /host/stateabove).
Residency truth is provider-first: provider_resident / provider_loaded
booleans in the source record outrank the runtime-lease booleans resident /
loaded, because a runtime can hold a lease on a model the provider has
already evicted. A loaded-looking state string (provider_loaded, loaded,
resident) can confirm residency, but a state string is never proof of
absence — with no boolean present and no loaded-like state, resident stays
null. details preserves the raw record for fields outside the schema.
POST /models/load accepts task (default text_generation), provider,
model, optional provider options, pin (default true), base_url,
timeout_s, and lock (default false) — with lock: true a successful
load is immediately locked against unload, and the lock outcome is reported
additively under lock in the response (a lock failure or a runtime without
lock support never turns the successful load into a failure). POST
/models/unload selects the runtime by runtime_id or by
task/provider/model. Both relay Runtime's host facade and return the
normalized residency response (operation, affected records, and in-band
ok=false errors instead of opaque failures). A load response carries
loaded_new (true when the call loaded the model, false when it was
already in memory) and the matching action (loaded / already_loaded).
A local image or video model loaded with task: "image_generation" (or
"video_generation") serves the following requests for that model on the loaded
pipeline, in the gateway process; a model that was not loaded runs each
request in an isolated worker process that loads it for that request.
One unload failure gets a real status code: when the target model is locked,
POST /models/unload answers HTTP 409 with the normalized refusal payload
as the body (ok: false, error: "model_locked", plus whatever detail the
runtime included), so clients can offer force-unload or point at
/models/unlock. Sending "force": true in the unload request unloads the
model despite the lock. Every other unload outcome stays in-band at 200.
Model locks and context estimates¶
POST /models/lockandPOST /models/unlock(admin) pin a resident model against unload and release that pin. The body selects the target like unload does:runtime_id, orprovider+model, with optionalbase_urlandtimeout_s. Lock requires provider-verified residency: a configured or merely-warm model refuses with anerror: "model_not_resident"payload (load it withlock: trueinstead); unlock always works, even for a since-evicted model, so locks are never stranded. Rows reportlockableso clients know whether a lock can work, andlockedso they can render the current state.GET /models/context_estimate?provider=&model=&context_length=(any authenticated principal) relays the Runtime host facade's context/KV memory estimate for a provider+model.providerandmodelare required;context_lengthis optional and must be >= 1 (schema-rejected with 422 otherwise). The estimate reports itsconfidencein-band —calibrated,estimated, orunknown— alongside facade fields such aspredicted_max_context(the context that fits beside the weights), the tri-statefits_weights/fits_requested_contextsplit,budget_bytes,est_kv_bytes, andnotes(which state the budget basis and reserve). The estimate is advisory only — no load path gates on it.
Like the other host-facade relays, these routes never 500 on capability gaps:
a runtime without the method answers 200 with ok: false,
available: false, and code = "model_residency_unavailable" (lock/unlock)
or code = "context_estimate_unavailable" (estimate); facade exceptions use
the matching *_error codes.
Session prompt-cache enumeration¶
GET /api/gateway/sessions/prompt_cache?session_id=<optional>POST /api/gateway/sessions/{session_id}/prompt_cache/clear_all(admin)
The list route enumerates the prompt caches the runtime actually minted. Each
cache row carries the provider/model/runtime identity, byte and token counts,
and stamped attribution metadata (session_id, run_id, workflow_id,
node_id). Omit session_id to list every session's caches. This
enumeration lane is the recommended way to observe and reclaim session cache
state: unlike the identity-derived session lifecycle endpoints described
under the prompt-cache control plane below, it cannot miss caches whose keys
the gateway never derived.
clear_all unloads every runtime-minted cache for one session in a single
call. It requires an admin principal because it accepts any session id and
clears real provider cache state; the identity-derived, caller-scoped session
lifecycle endpoints remain user-level.
When the runtime facade does not expose enumeration, both routes answer 200
with ok=false, available=false, and code="session_caches_unavailable"
(facade errors use code="session_caches_error"); the list route always
carries a caches array and clear_all always carries cleared and count.
Discovery descriptors¶
GET /discovery/capabilities advertises this surface under
capabilities.contracts.common:
model_residency:endpoints(loaded,load,unload,lock,unlock,context_estimate), the per-task support map,row_schema = "model_residency_row_v1", andmodality_ui— the canonical modality color map ({version: 1, colors: {...}}, one{color, label}entry per residency task plus anunknownfallback) so every client renders the same modality palette instead of hardcoding its own. It is a rendering contract, not a runtime capability, so it is served even when the runtime facade is absent.host_state:endpoints(state,memory,gpu) plusmemory_available. The state route itself always answers; per-section truth lives in the payload'sdegradedlist.session_caches:endpoints(list,clear_all) plusavailable, reflecting whether the runtime facade supports cache enumeration.
Evidence: src/abstractgateway/routes/gateway.py (host_state,
host_memory_metrics, model_residency_loaded, model_residency_lock,
model_context_estimate, session_prompt_caches_list) and
src/abstractgateway/security/authorization.py (route-family policy).
About (GET /api/gateway/about)¶
Public (no sign-in): which versions this gateway runs, for About screens.
{"abstractframework": "0.5.0", "abstractgateway": "0.7.1",
"packages": {"abstractcore": "2.18.0", "abstractruntime": "0.7.0", "abstractskill": "0.3.0"}}
abstractframework is the AbstractFramework release the installer recorded
for this gateway (FRAMEWORK_VERSION in the data directory's bootstrap.env)
when there is one, else the version of the abstractframework package
installed beside the gateway, else null (not installed). The installer puts
only abstractgateway[...] in the gateway's environment, so its record is
the one that names the release. Versions only: no paths, host names or settings.
Host control (pause, desktop tray, restart, update)¶
The process's own controls — the surface behind the desktop tray icon and the console's Gateway card (see tray.md). Reads are available to any authenticated principal; writes require an admin principal.
GET /api/gateway/host/runner— execution state:
{"ok": true, "paused": true, "paused_at": "2026-09-05T06:38:26+00:00", "paused_by": "default/admin", "reason": "meeting",
"inflight_ticks": 0, "scope": "workflow runner", "runner_in_process": true, "step_gate_supported": true,
"runners": [{"status": "paused", "...": "..."}], "degraded": false,
"capabilities": {"restart": true, "shutdown": true, "reason": null, "update_job_running": false}}
POST /api/gateway/host/pause(body{"reason": "..."}optional) andPOST /api/gateway/host/resume— both answer the payload above.GET /api/gateway/host/metrics/live—{gpu, memory, runner}in one call (1 s caches);gpu/memorycarry the same in-bandsupportedshape as/host/metrics/gpuand/host/metrics/memory.GET /api/gateway/host/runs?limit=25&window_hours=24— the runs on this host across every data plane (admin;/runsanswers only for the calling principal's plane).{ok, items: [{run_id, workflow_id, label, status, activity, role, created_at, updated_at, ledger_len, plane, started_epoch, updated_epoch, observer_path}], count, active_count, has_more, planes, skipped_entity_planes?, warnings?}. Items are TURN ROOTS (the runtime'sis_turn_root: parent-less runs that are not automation controllers, plus automation occurrences), in this order: every active root first —activity: "running"(the root or any run below it is running), then"waiting"(waiting for a person or an event, updated inside the window) — then roots that finished insidewindow_hoursof their last update (activity: "done"), newest first. Active rows are never windowed and never cut bylimit(active_countcounts them);limitbounds the finished rows andhas_moresays more finished ones exist.labeldecodes a catalog workflow's internal id (__catalog__v2__…<base64>) to the name an operator uses;observer_pathis the run's Observer page (/apps/observer/#run/<id>, open it signed in throughPOST /api/gateway/apps/observer/open {path: "/#run/<id>"}). The gateway's own bookkeeping runs (__-prefixed, but never a catalog id) are excluded.GET /api/gateway/host/tray—{dependencies_installed, install_hint, decision: {start, reason, hint}, supervisor: {running, ready, pid, exit_code, failure, log_path}, can_control}. There is no setting: the icon is shown whenever this process and this desktop can hold it.POST /api/gateway/host/tray/show— retry the helper now (409 when this process cannot). Nohidecounterpart, by design.POST /api/gateway/host/restart,POST /api/gateway/host/shutdown—{"ok": true, "restart": true, "requested_by": "...", "reason": "..."}; 409 with a plain reason when unsupported (--reload, embedded server, an update is installing).GET /api/gateway/host/start-at-login,PUT /api/gateway/host/start-at-login {enabled, replace_other?}— admin. Would this gateway start at the next login, and can it be changed from here:{enabled, state: on|off|broken|other, mechanism, mechanism_label, can_change, reason, summary, problems, location, platform, experimental[, other_data_dir]}. The mechanism is a LaunchAgent (macOS), a systemd user unit or, without a systemd user manager, a desktop autostart entry (Linux), or aHKCU\…\Runvalue (Windows).can_change: falsecarries the reason (for example a Linux server with neither a systemd user manager nor a desktop session).PUTregisters for the next login without starting a second gateway, or unregisters without stopping this one, and answers the read-back understart_at_login; 409 when it cannot change here or when another gateway's registration would be replaced withoutreplace_other: true; 500 when the change does not read back. The same switch as the tray's Start at login andabstractgateway service enable|disable.GET /api/gateway/host/update,POST /api/gateway/host/update/check,POST /api/gateway/host/update/start(admin) — all three answer{current, install: {kind, path, upgradable, reason, command, display_command, extras, data_dir, framework_version}, check: {source, latest, update_available, offline, checked_at, error, release}, job: {state, command, log_tail, exit_code, error, message, restart_recommended, version_before, version_after, installer, framework_before, framework_after, changes, other_changes}, restart_pending, update}.install.kindisinstallerfor an AbstractFramework installer install (check.source: framework-release,check.release: {version, gateway_version, installed, commit, manifest_url, installer: {url, commit_url, sha256, size}}); other kinds compare with PyPI (source: pypi).updateis the rendering every client shows:{status, line, hint, offer, action: {label, confirm, command, source, installer_sha256} | null, checked_at},statusone ofnot_checked,available,not_possible,up_to_date,offline,error,running,installed,no_change,failed.starttakes{installer_sha256}(the action's; required for an installer install, optional otherwise) and answers 409 when the install cannot be upgraded in place, the sha256 is missing ("check again first") or differs from the last check's, or a job is already running. The job is killed after 30 minutes even when it prints nothing. A version the check cannot compare is anerror, neverup_to_date.
GET /api/health adds "paused": true while paused; status stays
"healthy".
Prompt-cache control plane (operator API)¶
The gateway exposes prompt-cache operator endpoints under /api/gateway/prompt_cache/*.
Provider prompt-cache controls affect process-local or remote provider state
and require an admin principal in hosted user-auth mode.
Core endpoints:
GET /api/gateway/prompt_cache/capabilities?provider=...&model=...GET /api/gateway/prompt_cache/stats?provider=...&model=...POST /api/gateway/prompt_cache/setPOST /api/gateway/prompt_cache/updatePOST /api/gateway/prompt_cache/forkPOST /api/gateway/prompt_cache/clearPOST /api/gateway/prompt_cache/prepare_modules
Behavior:
- These routes use the runtime's AbstractCore prompt-cache client contract rather than directly depending on provider-instance access.
- In local mode they delegate to the in-process provider.
- In remote/hybrid mode they follow whatever
/acore/prompt_cache/*surface the configured AbstractCore server exposes. - All core prompt-cache responses include
operationandcapabilities, with structured unsupported/error cases (code="prompt_cache_unsupported"/code="prompt_cache_error"/code="prompt_cache_unavailable"). - These endpoints remain provider/model controls, not a Gateway-owned CachedSession persistence system.
Session lifecycle endpoints:
GET /api/gateway/sessions/{session_id}/prompt_cache/statusPOST /api/gateway/sessions/{session_id}/prompt_cache/preparePOST /api/gateway/sessions/{session_id}/prompt_cache/rebuildPOST /api/gateway/sessions/{session_id}/prompt_cache/clear
These routes derive a deterministic bounded namespace/key from session_id,
bundle_id, bundle_version, flow_id, provider, model, optional
template_id, and version. The private hash also includes the authenticated
principal scope, so two hosted users using the same session id/provider/model do
not collide in a shared provider control plane; the returned identity remains
portable app-level data and does not expose that private scope. These routes
expose three honest modes:
unsupported: provider/model does not expose prompt-cache support; responses includesupported=false,ok=false, and capabilities.keyed: gateway returns a stableruntime_hint/prompt_cache_keyfor Runtime/Core injection, but does not claim module preparation occurred.local_control_plane: gateway uses supported provider operations such asprepare_modules,fork,set,clear, andstats.
status is read-only. prepare accepts optional modules (system_prompt,
workflow_instructions, tools, pinned_attachments) and returns either
provider operation results or a key hint. rebuild is clear-plus-prepare for
providers that expose clear controls.
These identity-derived endpoints only see caches whose keys the gateway derived. To enumerate or bulk-clear the caches the runtime actually minted for a session, use the recommended session prompt-cache enumeration lane.
Durable bloc exact-reuse endpoints:
POST /api/gateway/blocs/upsert_textGET /api/gateway/blocs/recordGET /api/gateway/blocsPOST /api/gateway/blocs/deleteGET /api/gateway/blocs/kv/manifestGET /api/gateway/blocs/kv/listPOST /api/gateway/blocs/kv/ensurePOST /api/gateway/blocs/kv/loadPOST /api/gateway/blocs/kv/deletePOST /api/gateway/blocs/kv/prune
These routes are the primary app-facing durable prompt-cache path:
- create or identify a durable text bloc;
- ensure or load a KV artifact for a target local provider/model;
- use the returned
prompt_cache_bindingin later Runtime-backed generation; - list/delete/prune artifacts without reaching into provider-private cache state.
They delegate through Runtime's public AbstractCore host facade rather than
proxying Core directly. They are operator-style host controls, so the routes
themselves are not ledgered run execution; the ledgered exact-reuse path is the
later LLM_CALL.params.prompt_cache_binding used inside real Runtime runs.
Host-local prompt-cache export/import admin aliases:
GET /api/gateway/prompt_cache/savedPOST /api/gateway/prompt_cache/savePOST /api/gateway/prompt_cache/load
These remain explicitly local/operator-oriented:
- the route paths are compatibility aliases, but the implementation delegates to Runtime's public host facade:
saved->list_prompt_cache_exports(...)save->prompt_cache_export(...)load->prompt_cache_import(...)- local bundle/file runtimes store these exports under the Gateway data dir at
prompt_cache_exports/ - remote and hybrid runtimes return
code=prompt_cache_local_only - response payloads follow Runtime's host-local export/import contract, including
operation,local_only,artifact_*,capabilities, andprovider_response
Accounts and activity¶
The Accounts page (admin): users and entities in one list.
| Route | Purpose |
|---|---|
GET /admin/accounts |
{accounts: [{id, tenant_id, kind: user \| entity, role: admin \| user \| entity, own, email_address, mailbox{state: connected \| receive_only \| not_connected \| paused \| unavailable, address, provider, reason}, runtime_id, active, entity_state, actions{email, logs, workspace, preferences, rotate, manage, delete, suspend: {available, reason}}}]}, sorted admins, users, entities, then id; entities_warning when the entity list could not be read |
PUT /admin/accounts/{id}/active |
{active} → the updated row. Users: false = deactivated (signed out, cannot sign in); 409 {message} for your own account ("You can't deactivate your own account.") or the last active admin. Entities: false = suspended (entity state paused, its door credential off, an open visit closed); true = resumed (the state it had before is restored, stored in <data_dir>/auth/entity_suspended.json) |
GET /admin/accounts/{id}/activity |
?limit=100&kind=sign_in,run,… (admin) |
GET /me/activity |
the same for the signed-in account |
GET /me/accounts |
any signed-in account: {accounts: [rows], scope: "own"}, the same row shape — your own row plus one row per entity you created; actions only an admin can take are unavailable with the reason |
GET /me/accounts/{id}/activity |
the activity of your own account or of an entity you created; any other id answers 404 |
Who sees which account¶
An admin sees every user and every entity (GET /admin/accounts; non-admins get 403 there). Anyone
else sees only themself and the entities they created. Entity rows carry created_by
({tenant_id, user_id} of the account whose POST /entities created it, or null).
This is enforced on every entity route, not only on the Accounts page:
POST /entitieswritescreated_byinto the new home'smanifest.json. Entities created before this field existed have none and are visible to admins only: no creator is guessed and no existing manifest is rewritten.GET /entitieslists only the entities the caller may see.- Every
/entities/{name}/…route (inspect, card, state, chat, visit, summon, workspace, tool policy, replay, …) checks first. An entity you may not see answers exactly like one that does not exist: 404 with the same sentence, so names cannot be probed this way. POST /entities/meets/openneeds both entities visible;/entities/meets/{id}answers 404 for a meet with an entity you may not see.- Creating an entity under a name another account already holds answers 409 ("That name is taken
…"): entity names are unique per gateway, so this is the one place a name's existence shows. This
holds across runtimes: an entity living in another user's runtime, or a user account with that
name, also answers 409 (
POST /entities/{name}/validatereports the same error). - Visibility is not management: entity writes that are admin-only (state, tool policy, prompt, substrate, …) stay admin-only for the entities you created.
Without user accounts (the single-operator gateway) every entity is visible, as before.
The email address and mailbox of a row come from the resolver GET /me/email uses. An action that
cannot apply says why in reason: an entity has no mailbox ("Entities can't have their own mailbox
yet: mailboxes belong to a user's runtime." — entity runs use their own runtime, without the host's
mailbox), no token to rotate (its credential is discarded at creation) and no delete ("An entity's
name is kept for life; suspend it instead.").
Activity answers {events: [{ts, kind: sign_in \| token \| run \| automation \| email \| account,
title, detail, run_id, observer_path, ok, ts_local}], source: "audit_log", oldest_ts, truncated, note}, newest
first. It reads <data_dir>/audit_log.jsonl and its rotated files backwards within a fixed read
budget (truncated: older entries were not read). Events come from two explicit tables: the request
lines (sign-ins, sign-out, runs started with their run_id, automation commands, account changes by
an admin) and the typed email events (mailbox connected, tested, notification sent or not, …). The
audit log records writes only: read-only requests (page views, token use on reads), mail received and
what agents send with their email tools are not in it, and note says so. observer_path opens the event in the Observer app:
/apps/observer/#run/<run_id> for a run, /apps/observer/#automations for an automation event; ts_local
is ts in the gateway's local time (ISO 8601 with its offset).
A run id is created by the run-start route, so it is not in the request path: POST /runs/start and
POST /runs/schedule write run {run_id, workflow, bundle_id, bundle_version, entrypoint, scheduled} on
their audit line (workflow is the entrypoint's name, else the bundle id). A run event's detail is that
workflow name. A run-start line written by a gateway older than 0.10.0 has no run id: its detail is "Run id not
recorded (before this version)" and it has no run_id and no observer_path (nothing is guessed from
times). A notification event's detail is its kind in words: "Approval needed", "Job failed", "Job
finished", "Automation result", "Automation failed", "Test notification"; any other kind is shown as
written.
Account preferences¶
The gateway keeps each account's client preferences, so every app and every device of that account see the same choice and no client keeps it. The gateway defines the defaults; an account may override them for itself.
| Route | Purpose |
|---|---|
GET /api/gateway/accounts/{account}/preferences |
{ok, account: "tenant:user", can_edit, preferences: {default_workflow: {<interface>: value \| null}}, declared: {default_workflow: {label, help}}, apps: [row]} |
PUT /api/gateway/accounts/{account}/preferences |
{"default_workflow": {<interface>: "bundle:flow" \| "catalog:bundle:flow" \| null}} → the GET answer |
{account} is me (the caller), name, tenant:name or an entity's slug. Who: anyone for their
own; an admin for any account; an entity's preferences by an admin or the entity's creator. Anyone
else gets 403 with "Only an admin or the account itself can read or change its preferences." (an entity:
"Only an admin or
Only keys the gateway declares are accepted. Today there is one, default_workflow: one entry per
app interface (abstractcode.agent.v1, abstractassistant.agent.v1). null follows the
gateway's per-app default, which is the admin setting agents.default_workflow.<interface>
(Workflows → Default workflow per app, configuration.md); it
is never copied into the account. A chosen workflow is stored without a version, so it runs its
latest published version. The PUT replaces the named interfaces and keeps the others. It is
refused with 400 {"detail": {"reason": "preference_refused", "message": <sentence>, "key": <key>}}
for an unknown key ("Unknown preference 'theme': this gateway declares default_workflow."), an
interface that is not an app's, or a workflow the account may not run for that app ("…refused:
'helper@1.0.0:agent' declares abstractassistant.agent.v1, not abstractcode.agent.v1"). A refused PUT
changes nothing. The change is audited on the target account's activity ("Preferences changed").
Each apps row: {interface, label, app, help, value, state, reason, gateway_default,
gateway_default_label, effective, choices}.
| Field | Meaning |
|---|---|
value |
the account's choice, null = the gateway default |
state |
default (null), set (a choice that runs), broken (a choice that no longer runs; reason says why and what to do, e.g. "coder-two:agent no longer runs for alice: … Pick another workflow or Gateway default.") |
gateway_default |
{available, name, value, workflow_id, reason}: what the admin's per-app default resolves to now |
gateway_default_label |
"Gateway default ( |
effective |
what a run of this app starts for this account: {source: account \| gateway, available, name, value, workflow_id, bundle_id, bundle_version, flow_id, registry_scope, reason} |
choices |
the workflows the account may run for this app: [{value, label, name, workflow_id, bundle_id, bundle_version, flow_id, registry_scope}], latest versions, never one an admin made unavailable to users or archived; label is shown verbatim |
For me the choices are the caller's own (shared workflows available to them and their own
workflows). For another account they are the gateway's shared workflows that account may run:
an admin is never offered a user's private workflows.
Runs: an app starts the account's choice as an explicit workflow, or flow_id: "@default" with
its interface when the value is null. @default keeps meaning the admin's per-app default.
Accounts rows (GET /admin/accounts, GET /me/accounts) carry actions.preferences
{available, reason} (every user and entity that is not archived).
Email¶
Per-user email (email.md): every route acts on the caller's own account, resolved from the authenticated principal (never from a path or body id); entities are refused (403). No response ever carries a password, token or OAuth client secret.
The email address is the user's own address (sign-in codes, notifications, the default allowed recipient; no password); the mailbox is the connection the user makes so their agents and automations can read and send mail as them.
Your email address and mailbox (/api/gateway/me/..., every signed-in human):
| Route | Purpose |
|---|---|
GET /me/email |
settings and status (mailbox{state, address, provider, reason} as in /admin/accounts): configured, address, imap, smtp, auth_kind, secret_set, policy, limits (+ usage), status{last_test, last_ok, last_error{code, cause, fix}}, watcher{state, last_poll, cursor, received}, admin_enabled, effective_enabled; for the account page: email_address (your email address as stored), registered_address ("self" for runs: that address, else the mailbox's own), email_available, notifications{job_failed, approval_needed} + notifications_unavailable_reason, agent_tools{on, available, unavailable_reason, active}, oauth_providers[{id, available, reason}]; send_capable (false for a mailbox stored without its SMTP leg — mailbox.state: receive_only) |
POST /me/email/discover |
{address} → {address, domain, found, source, provider, imap, smtp, username, tried, defaults} (known providers, autoconfig, ISPDB, SRV, MX); defaults is what the form pre-fills: {imap{host, port, security}, smtp{…}, login, source: discovered \| standard, provider, message} (AbstractCore server_defaults: the discovered servers, else imap.<domain> 993 SSL and smtp.<domain> 465 SSL, "Standard settings for |
PUT /me/email |
connect (save + test): {address, password, username?, display_name?, imap?{host, port, security, folder}, smtp?{host, port, security}, test}; username defaults to the discovered login, else the address; display_name to the stored name, else the address's local part; connecting sets your email address when it is empty (audited email.address_changed, reason mailbox_connected); without imap and smtp the servers are discovered (discovery{source, provider, tried} in the answer; none found = 400 email_discovery_failed with tried); signs in to both servers first (unless test: false), stores nothing on a failure and names the failing step (detail.step: imap | smtp, detail.message); a leg left out is filled with the domain's standard server (or the discovered one) and the test signs in to BOTH, so a connected mailbox can always send (an unreachable outgoing server is detail.step: smtp, nothing stored); tools_reloaded: the caller's workflows were reloaded so agents' toolsets follow |
PUT /me/email/address |
{address} — your email address, stored on your user record ("" clears it; 400 email_invalid_settings for anything but one plain address) |
PUT /me/email/notifications |
{job_failed?, approval_needed?} — the two notification switches (both on by default); the earlier {email: {...}} body is accepted and mapped; answers like GET /me/email |
POST /me/email/test |
per-leg result {imap, smtp, ok, message}; message: "Test passed: signed in to imap.x and smtp.x." or the failing step and its cause ("Sign-in refused by imap.x — check the password. (…)") |
DELETE /me/email |
disconnect: credentials and cursor deleted (policy and limits kept); tools_reloaded |
PUT /me/email/policy |
{mode: "allowlist" \| "denylist", always_allow?: [address \| domain], always_deny?: [address \| domain]}; a list given replaces that list. Precedence: own address → allowed; Always denied → refused; Always allowed → allowed; else the mode (allowlist = Only the Allowed list → refused, denylist = Anyone not on the Denied list → allowed). A domain covers its subdomains. The older {mode, entries} is accepted (entries = the mode's list) |
POST /me/email/policy/check |
{to?, cc?, bcc?, addresses?} (addresses = To) → per-recipient verdicts with source (self, always_deny, always_allow, mode) |
PUT /me/email/limits |
{per_hour, per_day} |
PUT /me/email/folder |
{folder} — the folder your mailbox is read from (empty = INBOX); the connection is kept and nothing is tested; 404 email_not_configured without a mailbox |
PUT /me/email/enabled |
{enabled} — "Use this mailbox" (off keeps the settings; stops watching, sending and notifications); tools_reloaded |
PUT /me/email/agent-tools |
{enabled} — your agents' email tools (default off; 409 email_disabled while "Agent email tools for users" is off for you; active only with a connected, allowed mailbox); reloads your workflows so toolsets follow (tools_reloaded) |
GET /me/email/oauth/clients |
which providers have a gateway OAuth client (no secrets) |
POST /me/email/oauth/start |
{address, provider: google \| microsoft, client_id?, client_secret?, tenant?, flow?: device \| loopback} → device code (user_code, verification_uri) or authorization_url. token_endpoint, authorization_endpoint, device_authorization_endpoint and scopes are accepted only from an administrator with provider: "custom" (own client id); otherwise 403 email_oauth_override_refused |
POST /me/email/oauth/poll |
{flow_id} → {pending: true} or the connected account |
POST /me/email/oauth/finish |
{flow_id, wait_s} — waits up to 60 s for the approval; tools_reloaded once connected |
POST /me/email/oauth/cancel |
{flow_id} |
GET /me/notifications |
events and choices, channel availability, outbox summary |
PUT /me/notifications |
{email: {job_failed, approval_needed: bool}}; the earlier kinds are accepted (automation_failed counts for job_failed; automation_result and job_finished are per-automation / per-run options now) |
POST /me/notifications/test |
sends one test notification now → {ok, sent, reason_code, message, limit, state, error?}: reason_code is null (sent) or no_mailbox, mailbox_paused, rate_limited, queued_behind (+ queued_behind: how many), send_failed; message is the sentence to show ("Sent to x@y.", "Not sent: hourly limit reached (20 of 20 this hour) — resets at 14:05.", in the gateway's local time); limit {window: hour \| day, limit, used, resets_at} when a send limit held it back |
Administrators (status and the switch only; administrators never read mail):
| Route | Purpose |
|---|---|
GET /admin/users |
every row carries email_address and mailbox{state: connected \| receive_only \| not_connected \| paused \| unavailable, address, provider, reason} from the same resolver as GET /me/email (so your own row matches your card; email stays the raw record field); each human row also carries email_account: {configured, address, state, admin_enabled, agent_tools_available, capabilities} (capabilities: {email, email_agent_tools} as {value, source: user \| gateway \| built-in}; user = a per-user override) |
GET /admin/users/{user_id}/email |
configured, address, auth_kind, user_enabled, admin_enabled, effective_enabled, status (last test / last error), capabilities ({value, source: user \| gateway \| built-in}), agent_tools (available, user_enabled, active), watcher, state |
PUT /admin/users/{user_id}/email |
{enabled?, agent_tools?, inherit?: ["email", "email_agent_tools"]} — per-user capabilities; enabled: false = no watcher, no sending, no notifications (settings kept); agent_tools = Agent email tools available |
GET /admin/email/capabilities |
capabilities[{id, label, description, per_user, advanced, default, built_in_default}]: email "Mailboxes for users" (on), email_agent_tools "Agent email tools for users" (on) and email_recovery "Sign-in by email" (on) |
PUT /admin/email/capabilities |
{email?, email_agent_tools?, email_recovery?, reset?: [...]}; tools_reloaded: how many built hosts were rebuilt so every user's agents follow the change |
GET /admin/email/oauth-clients |
bring-your-own OAuth clients: client_id, client_secret_set, tenant per provider |
PUT /admin/email/oauth-clients/{provider} |
{client_id, client_secret?, tenant?}; an empty client_id removes the provider's client; omitting client_secret keeps the stored one for the same id |
Sign-in page (public):
| Route | Purpose |
|---|---|
GET /session/recovery |
{available} — true when sign-in by email is on (email_recovery, default on) and at least one account of this gateway has email |
POST /session/recovery/request |
{user_id, tenant_id?, purpose?: sign_in (default) \| reset_token} → {sent: true, to: "l•••@•••", expires_in_s, message}, {sent: false, reason_code: "no_email_address", message} (also for an unknown account), {sent: false, reason_code: "no_mailbox", message} (an address but no mailbox to send from), {sent: false, reason_code: "send_failed", message} (the mail server refused or could not be reached; the request waits up to 12 s for the real outcome) or {sent: false, reason_code: "too_many_requests", retry_after_s, message}; 404 recovery_off when sign-in by email is off (email.md) |
POST /session/recovery/redeem |
{user_id, tenant_id?, purpose?, code, remember?} → a browser session; reset_token also returns the new token once. A wrong, expired or used code answers 401 recovery_code_refused |
Errors carry {"detail": {"reason_code", "message", "cause", "fix", "retryable"}}: 400 invalid
settings or a policy refusal (email_policy_refused), 403 entity or a refused OAuth override (email_oauth_override_refused), 404 no account
(email_not_configured), 409 turned off (email_disabled) or credentials missing, 422 the mail
server refused (email_auth_failed, email_tls_failed, email_unreachable, …), 429 send limit
(email_rate_limited).
curl -sS -H "$AUTH" "$BASE_URL/api/gateway/me/email"
curl -sS -X POST -H "$AUTH" "$BASE_URL/api/gateway/me/email/test"
curl -sS -X POST -H "$AUTH" "$BASE_URL/api/gateway/me/notifications/test"
Runs: _runtime.notify = {"on": ["finished", "failed"], "channels": ["email"]} in a run's input asks
for a notice when it ends (unknown values are dropped). _runtime.email_account and
_runtime.email_allowed_recipients are set by the gateway; client-supplied values are removed.
Deprecated aliases (admin only, the calling administrator's own account; removed in a later minor
release): GET /email/accounts, GET /email/messages (mailbox, since, status, limit),
GET /email/messages/{uid} (whole body, content_trust: "untrusted"), POST /email/send
({to, cc?, bcc?, subject, body_text?, body_html?}; the recipient policy and limits apply).
Troubleshooting and common questions: faq.md.