Skip to content

Troubleshooting

This page is for symptom-oriented fixes. For concepts and limits, see faq.md; for setup and examples, see getting-started.md.

Importing abstractruntime.integrations.abstractcore fails

Symptom: - Importing the AbstractCore integration raises an ImportError about the required AbstractCore version or missing optional dependencies.

Checks:

python -m pip show AbstractRuntime abstractcore
python -m pip install -U abstractruntime

Fix: - Install or upgrade the base Runtime package. LLM/tools integration, common remote-light multimodal dependencies, and the MCP worker entry point are part of the base install. - The current AbstractCore integration expects abstractcore>=2.20.2.

Verify:

python -c "import abstractruntime.integrations.abstractcore as ac; print(ac.__all__[:5])"

A run is waiting and does not continue

Symptom: - Runtime.tick(...) returns status=waiting, and the run stays paused.

Likely causes: - The run is waiting for ASK_USER, WAIT_EVENT, passthrough tools, or tool approval. - WAIT_UNTIL needs a driver loop, or a host must call tick(...) again after the due time.

Fix: - For user/event/tool waits, resume with the exact wait_key from state.waiting.wait_key. - For time-based waits, use create_scheduled_runtime() or another host driver that periodically ticks due runs.

Verify:

state = rt.get_state(run_id)
print(state.status, state.waiting.wait_key if state.waiting else None)

Generated media is missing or too large for run state

Symptom: - Image, voice, music, or transcription outputs fail, or binary content is not present in RunState.vars.

Likely causes: - Generated binary outputs require a runtime ArtifactStore. - Runtime stores generated bytes by artifact reference instead of embedding raw bytes in checkpoints or ledger records.

Fix: - Construct the runtime with an artifact store such as InMemoryArtifactStore or FileArtifactStore. - Read artifact_id / artifact_ref from the durable result and load the artifact from the configured store.

Docs: - api.md#artifacts-store-by-reference - integrations/abstractcore.md#multimodal-generation

Remote media calls use the wrong model or fail with input-media errors

Symptom: - Remote image/TTS/STT calls do not use the expected media model. - Remote image generation fails when media is supplied. - Remote transcription rejects media that is not a local file or artifact-backed temporary file.

Fix: - Put endpoint-specific routing in the output selector, not in the chat model unless you intend to route chat. - Use output.task="image_edit" for image edits with one source image and optional mask. - Use exactly one audio media item for remote STT, and make sure it resolves to a local path or artifact-backed file.

Docs: - integrations/abstractcore.md#multimodal-generation

Local media residency returns model_residency_unsupported

Symptom: - A local MODEL_RESIDENCY load for image_generation, image_upscale, video_generation, text_to_video, image_to_video, tts, stt, or music_generation returns ok=false with code="model_residency_unsupported".

Meaning: - Runtime is reporting that the current local execution topology cannot truthfully keep that media backend resident for later reuse.

Fix: - Use a configured long-lived AbstractCore server for media residency, then let Runtime relay /acore/models/*. - Keep local media warmup optional (required=false) if it is only an optimization. - Do not infer loaded state from model defaults, catalogs, downloaded weights, or provider names.

Docs: - integrations/abstractcore.md#prompt-cache-control-plane-and-durable-blocs - faq.md#why-can-local-media-residency-return-okfalse-without-failing-the-run

Prompt-cache behavior is not reused for generated media

Symptom: - A workflow-level prompt-cache flag creates session cache keys for text/chat calls, but not for image, voice, music, or transcription output selectors.

Meaning: - Runtime only auto-derives session prompt-cache keys for text/chat calls. Non-text output selectors may still carry an explicit prompt_cache_binding, but Runtime does not invent one.

Fix: - Use params.prompt_cache_binding for durable exact text/chat prefix reuse. - Treat generated media calls as separate capability executions unless the selected AbstractCore backend documents a task-specific cache contract.

Docs: - integrations/abstractcore.md#prompt-cache-control-plane-and-durable-blocs - faq.md#where-should-cached-session-or-prompt-cache-state-live

Live replies do not stream

Symptom: - A host registered a sink with Runtime.set_live_delta_sink(...), but an answer arrives only when the call ends, or the sink receives only an llm.delta_end.

Checks: - The run must set _runtime.stream to the boolean True; Runtime.start refuses any other value with a ValueError. - Read the call's llm.delta_end: reason: "unavailable" carries a detail, and the LLM_CALL ledger record carries the same value as metadata._runtime_observability.stream_unavailable.

Fix, by detail: - remote_core: remote mode does not stream; use a local or multi-local runtime for live text. - node_stream_off: the node sets params.stream: False; remove it to stream that node. - structured_output: structured and media-output calls never stream; this is expected. - usage_unavailable: the provider's server cannot report token usage in streams; configure the server to accept stream_options, or accept non-streamed answers for that model. - tool_envelope_holdback: the whole answer was a tool call or a hidden channel; this is expected. - sink_error: your sink raised. Keep the sink fast and non-raising (put the event on a queue and return).

Verify: - Start a run with vars={"_runtime": {"stream": True}} and check that the sink receives llm.delta events before the llm.delta_end with reason: "completed".

Docs: - integrations/abstractcore.md#live-token-streaming

The previous model stays in memory after a default switch

Symptom: - After set_default_provider_model(...), the previous MLX, HuggingFace or embedding model still appears in list_model_residency and memory does not drop.

Checks: - Read list_model_residency diagnostics. pending_ejects lists a model waiting for its running call to end. last_switch_ejects shows, per model, whether it was unloaded, kept because another owner still uses or locked it (skipped with a reason), or failed (ok: false with the remaining holders).

Fix: - A pending model is unloaded when its call ends; no action is needed. - A model kept for another owner stays until that owner releases it. Unlock or unload it there. An explicit unload_model_residency(runtime_id=...) frees the model from every holder in the process, including owners that still use it, so use it only when you intend that. - A skipped reason that names the claim registry means the installed AbstractCore predates it; upgrade AbstractCore.

Docs: - integrations/abstractcore.md#unloading-and-switching-models

An automation does not fire

Symptom: - get_automation(...) shows the automation, but no new occurrence appears in list_occurrences(...).

Likely causes: - Nothing drives the controller. Without a host run loop, drive_automation returns when the controller parks; a later tick needs the controller to be ticked again after its deadline. - The automation is paused, archived, completed (its schedule is exhausted: count reached, until passed, or a one-shot schedule already fired) or failed. - An occurrence is still running or waiting on a person, or is in retry backoff. Occurrences run one at a time. - The trigger is manual@1, which only fires on automation.run_now.

Checks:

from abstractruntime.automations import get_automation, list_occurrences, pending_waits

info = get_automation(runtime.run_store, automation_id)
print(info["status"], info["next_fire_at"], info["state"]["pending_occurrence"])
print(pending_waits(runtime.run_store, automation_id))

Fix: - Standalone hosts: tick the controller after next_fire_at, then drive it: runtime.tick(workflow=controller_workflow_spec(), run_id=automation_id) and drive_automation(runtime, automation_id). - Answer a pending wait with Runtime.resume(...) and the payload of its kind, or send automation.stop_current. - Send automation.resume for a paused automation. It re-arms at the next tick; automation.run_now fires once immediately.

Verify: - list_occurrences(...) shows a new item, and next_fire_at moves to the following tick.

Docs: - automations.md#the-controller, automations.md#commands

A start fails with identity_conflict

Symptom: - Runtime.start(..., run_id=...) raises RunIdentityConflict, or create_automation, start_discussion or apply_automation_command report identity_conflict.

Likely causes: - The id (or request_id, or command_id) was already used for a different request. Ids are idempotency keys: the same request returns the existing run or result, a different one is refused.

Fix: - Use a new request_id / command_id for a new request, or resend exactly the original request to get the existing result.

Docs: - api.md#runtime-start--tick--resume, automations.md#creating-an-automation, automations.md#commands

A start fails with SessionAttributionError

Symptom: - Runtime.start(..., session_id=...) raises SessionAttributionError (reason_code = "session_attribution_failed").

Likely causes: - The run store has no run index (neither session_kinds nor list_run_index), so the runtime cannot tell whether the session is a discussion. - The session is a discussion whose root cannot be validated: its runs name different roots, or the root is missing, has a parent, lives in another session, carries no seed, or has no workspace_root.

Fix: - Use one of the built-in run stores (SQLite, JSON files, in-memory, optionally wrapped by the offloading store), or add list_run_index / session_kinds to your store. - For a damaged discussion session, start a new discussion with start_discussion(...) instead of adding turns to it.

Docs: - automations.md#discussions, automations.md#storage-guarantees

Growing context or a discussion fails with history_unavailable

Symptom: - An automation's admission fails, or start_discussion raises SessionHistoryError (reason_code = "history_unavailable").

Likely causes: - The run store has no run index; strict history needs one. - A discussion's seed is missing, malformed or offloaded to an artifact the artifact store cannot load. - through_occurrence names an occurrence the session does not hold.

Fix: - Use a store with a run index and pass the runtime's artifact store to history reads. - Check the occurrence number against list_occurrences(...).

Docs: - automations.md#context-independent-or-growing, automations.md#runs-sessions-and-history

A tool is refused because the workspace is read-only

Symptom: - A tool or VisualFlow node fails with "is inside a read-only mount" or "this workspace is read-only", or a run fails because a read-only workspace_root does not exist.

Likely causes: - "inside a read-only mount": the run is a discussion (or lists the folder in _runtime.workspace_read_only_paths) and the tool tried to write into the automation's workspace. Reads and commands are allowed; file writes into the mount are refused. - "this workspace is read-only": the run was started with workspace_read_only: true. Tools classified write or exec, tools the runtime does not classify, and file-writing nodes are refused. A read-only workspace folder is never created.

Fix: - In a discussion, write into its own workspace (workspace_root) instead of the mount. - Otherwise continue the work in the automation itself or in an ordinary chat session, where the workspace is writable. - Check a tool's class with tool_effect_class(name) from abstractruntime.integrations.abstractcore.tool_effects.

Docs: - automations.md#read-only-mounts, automations.md#read-only-workspaces

MCP worker command is not found

Symptom: - abstractruntime-mcp-worker is not available on the command line.

Fix:

python -m pip install -U abstractruntime
abstractruntime-mcp-worker --help

Docs: - mcp-worker.md