Troubleshooting¶
This page is for symptom-oriented fixes. For concepts and limits, see faq.md; for setup and examples, see
getting-started.md.
Importing abstractruntime.integrations.abstractcore fails¶
Symptom:
- Importing the AbstractCore integration raises an ImportError about the required AbstractCore version or missing
optional dependencies.
Checks:
python -m pip show AbstractRuntime abstractcore
python -m pip install -U abstractruntime
Fix:
- Install or upgrade the base Runtime package. LLM/tools integration, common
remote-light multimodal dependencies, and the MCP worker entry point are part
of the base install.
- The current AbstractCore integration expects abstractcore>=2.20.2.
Verify:
python -c "import abstractruntime.integrations.abstractcore as ac; print(ac.__all__[:5])"
A run is waiting and does not continue¶
Symptom:
- Runtime.tick(...) returns status=waiting, and the run stays paused.
Likely causes:
- The run is waiting for ASK_USER, WAIT_EVENT, passthrough tools, or tool approval.
- WAIT_UNTIL needs a driver loop, or a host must call tick(...) again after the due time.
Fix:
- For user/event/tool waits, resume with the exact wait_key from state.waiting.wait_key.
- For time-based waits, use create_scheduled_runtime() or another host driver that periodically ticks due runs.
Verify:
state = rt.get_state(run_id)
print(state.status, state.waiting.wait_key if state.waiting else None)
Generated media is missing or too large for run state¶
Symptom:
- Image, voice, music, or transcription outputs fail, or binary content is not present in RunState.vars.
Likely causes:
- Generated binary outputs require a runtime ArtifactStore.
- Runtime stores generated bytes by artifact reference instead of embedding raw bytes in checkpoints or ledger records.
Fix:
- Construct the runtime with an artifact store such as InMemoryArtifactStore or FileArtifactStore.
- Read artifact_id / artifact_ref from the durable result and load the artifact from the configured store.
Docs:
- api.md#artifacts-store-by-reference
- integrations/abstractcore.md#multimodal-generation
Remote media calls use the wrong model or fail with input-media errors¶
Symptom: - Remote image/TTS/STT calls do not use the expected media model. - Remote image generation fails when media is supplied. - Remote transcription rejects media that is not a local file or artifact-backed temporary file.
Fix:
- Put endpoint-specific routing in the output selector, not in the chat model unless you intend to route chat.
- Use output.task="image_edit" for image edits with one source image and optional mask.
- Use exactly one audio media item for remote STT, and make sure it resolves to a local path or artifact-backed file.
Docs:
- integrations/abstractcore.md#multimodal-generation
Local media residency returns model_residency_unsupported¶
Symptom:
- A local MODEL_RESIDENCY load for image_generation, image_upscale, video_generation, text_to_video, image_to_video, tts, stt, or music_generation returns ok=false with
code="model_residency_unsupported".
Meaning: - Runtime is reporting that the current local execution topology cannot truthfully keep that media backend resident for later reuse.
Fix:
- Use a configured long-lived AbstractCore server for media residency, then let Runtime relay /acore/models/*.
- Keep local media warmup optional (required=false) if it is only an optimization.
- Do not infer loaded state from model defaults, catalogs, downloaded weights, or provider names.
Docs:
- integrations/abstractcore.md#prompt-cache-control-plane-and-durable-blocs
- faq.md#why-can-local-media-residency-return-okfalse-without-failing-the-run
Prompt-cache behavior is not reused for generated media¶
Symptom: - A workflow-level prompt-cache flag creates session cache keys for text/chat calls, but not for image, voice, music, or transcription output selectors.
Meaning:
- Runtime only auto-derives session prompt-cache keys for text/chat calls. Non-text output selectors may still carry an
explicit prompt_cache_binding, but Runtime does not invent one.
Fix:
- Use params.prompt_cache_binding for durable exact text/chat prefix reuse.
- Treat generated media calls as separate capability executions unless the selected AbstractCore backend documents a
task-specific cache contract.
Docs:
- integrations/abstractcore.md#prompt-cache-control-plane-and-durable-blocs
- faq.md#where-should-cached-session-or-prompt-cache-state-live
Live replies do not stream¶
Symptom:
- A host registered a sink with Runtime.set_live_delta_sink(...), but an answer arrives only when the call ends, or the
sink receives only an llm.delta_end.
Checks:
- The run must set _runtime.stream to the boolean True; Runtime.start refuses any other value with a ValueError.
- Read the call's llm.delta_end: reason: "unavailable" carries a detail, and the LLM_CALL ledger record carries
the same value as metadata._runtime_observability.stream_unavailable.
Fix, by detail:
- remote_core: remote mode does not stream; use a local or multi-local runtime for live text.
- node_stream_off: the node sets params.stream: False; remove it to stream that node.
- structured_output: structured and media-output calls never stream; this is expected.
- usage_unavailable: the provider's server cannot report token usage in streams; configure the server to accept
stream_options, or accept non-streamed answers for that model.
- tool_envelope_holdback: the whole answer was a tool call or a hidden channel; this is expected.
- sink_error: your sink raised. Keep the sink fast and non-raising (put the event on a queue and return).
Verify:
- Start a run with vars={"_runtime": {"stream": True}} and check that the sink receives llm.delta events before the
llm.delta_end with reason: "completed".
Docs:
- integrations/abstractcore.md#live-token-streaming
The previous model stays in memory after a default switch¶
Symptom:
- After set_default_provider_model(...), the previous MLX, HuggingFace or embedding model still appears in
list_model_residency and memory does not drop.
Checks:
- Read list_model_residency diagnostics. pending_ejects lists a model waiting for its running call to end.
last_switch_ejects shows, per model, whether it was unloaded, kept because another owner still uses or locked it
(skipped with a reason), or failed (ok: false with the remaining holders).
Fix:
- A pending model is unloaded when its call ends; no action is needed.
- A model kept for another owner stays until that owner releases it. Unlock or unload it there. An explicit
unload_model_residency(runtime_id=...) frees the model from every holder in the process, including owners that
still use it, so use it only when you intend that.
- A skipped reason that names the claim registry means the installed AbstractCore predates it; upgrade AbstractCore.
Docs:
- integrations/abstractcore.md#unloading-and-switching-models
An automation does not fire¶
Symptom:
- get_automation(...) shows the automation, but no new occurrence appears in list_occurrences(...).
Likely causes:
- Nothing drives the controller. Without a host run loop, drive_automation returns when the controller parks; a
later tick needs the controller to be ticked again after its deadline.
- The automation is paused, archived, completed (its schedule is exhausted: count reached, until passed, or
a one-shot schedule already fired) or failed.
- An occurrence is still running or waiting on a person, or is in retry backoff. Occurrences run one at a time.
- The trigger is manual@1, which only fires on automation.run_now.
Checks:
from abstractruntime.automations import get_automation, list_occurrences, pending_waits
info = get_automation(runtime.run_store, automation_id)
print(info["status"], info["next_fire_at"], info["state"]["pending_occurrence"])
print(pending_waits(runtime.run_store, automation_id))
Fix:
- Standalone hosts: tick the controller after next_fire_at, then drive it:
runtime.tick(workflow=controller_workflow_spec(), run_id=automation_id) and drive_automation(runtime, automation_id).
- Answer a pending wait with Runtime.resume(...) and the payload of its kind, or send
automation.stop_current.
- Send automation.resume for a paused automation. It re-arms at the next tick; automation.run_now fires once
immediately.
Verify:
- list_occurrences(...) shows a new item, and next_fire_at moves to the following tick.
Docs:
- automations.md#the-controller, automations.md#commands
A start fails with identity_conflict¶
Symptom:
- Runtime.start(..., run_id=...) raises RunIdentityConflict, or create_automation, start_discussion or
apply_automation_command report identity_conflict.
Likely causes:
- The id (or request_id, or command_id) was already used for a different request. Ids are idempotency keys: the
same request returns the existing run or result, a different one is refused.
Fix:
- Use a new request_id / command_id for a new request, or resend exactly the original request to get the
existing result.
Docs:
- api.md#runtime-start--tick--resume, automations.md#creating-an-automation, automations.md#commands
A start fails with SessionAttributionError¶
Symptom:
- Runtime.start(..., session_id=...) raises SessionAttributionError (reason_code = "session_attribution_failed").
Likely causes:
- The run store has no run index (neither session_kinds nor list_run_index), so the runtime cannot tell whether
the session is a discussion.
- The session is a discussion whose root cannot be validated: its runs name different roots, or the root is missing,
has a parent, lives in another session, carries no seed, or has no workspace_root.
Fix:
- Use one of the built-in run stores (SQLite, JSON files, in-memory, optionally wrapped by the offloading store), or
add list_run_index / session_kinds to your store.
- For a damaged discussion session, start a new discussion with start_discussion(...) instead of adding turns to it.
Docs:
- automations.md#discussions, automations.md#storage-guarantees
Growing context or a discussion fails with history_unavailable¶
Symptom:
- An automation's admission fails, or start_discussion raises SessionHistoryError (reason_code =
"history_unavailable").
Likely causes:
- The run store has no run index; strict history needs one.
- A discussion's seed is missing, malformed or offloaded to an artifact the artifact store cannot load.
- through_occurrence names an occurrence the session does not hold.
Fix:
- Use a store with a run index and pass the runtime's artifact store to history reads.
- Check the occurrence number against list_occurrences(...).
Docs:
- automations.md#context-independent-or-growing, automations.md#runs-sessions-and-history
A tool is refused because the workspace is read-only¶
Symptom:
- A tool or VisualFlow node fails with "is inside a read-only mount" or "this workspace is read-only", or a run
fails because a read-only workspace_root does not exist.
Likely causes:
- "inside a read-only mount": the run is a discussion (or lists the folder in _runtime.workspace_read_only_paths)
and the tool tried to write into the automation's workspace. Reads and commands are allowed; file writes into the
mount are refused.
- "this workspace is read-only": the run was started with workspace_read_only: true. Tools classified write or
exec, tools the runtime does not classify, and file-writing nodes are refused. A read-only workspace folder is
never created.
Fix:
- In a discussion, write into its own workspace (workspace_root) instead of the mount.
- Otherwise continue the work in the automation itself or in an ordinary chat session, where the workspace is
writable.
- Check a tool's class with tool_effect_class(name) from
abstractruntime.integrations.abstractcore.tool_effects.
Docs:
- automations.md#read-only-mounts, automations.md#read-only-workspaces
MCP worker command is not found¶
Symptom:
- abstractruntime-mcp-worker is not available on the command line.
Fix:
python -m pip install -U abstractruntime
abstractruntime-mcp-worker --help
Docs:
- mcp-worker.md