Autonomous mode
Autonomous is the gearbox's third gear: decompose → execute → verify → correct. Use it when a wrong answer is worse than a slower one — multi-document extraction, high-stakes analysis, anything the Critic should be allowed to reject.
v0.7.14 turned that gearbox into a bounded harness. The model still reasons. The framework now decides when to stop, what the Critic is allowed to see, and whether every declared document was actually processed.
Copy-paste first: Autonomous workflow. This page is the full guide — why the changes existed, what improved, and how to use the new APIs.
When to use it
| Use Autonomous when… | Stay on Standard when… |
|---|---|
| The job is about a known set of documents | One lookup + one answer |
| You need an independent Critic / Refiner | Latency matters more than verification |
| A missed file is a wrong deliverable | Coverage is "best effort" |
You want a typed response_format that the loop must satisfy |
Free-form prose is the product |
Standard and Autonomous share the same tool loop. Loop guards, termination_reason, diagnostics, and the structured-output contract apply to both. Decomposition, Critic/Refiner, coverage follow-up, and COMPLEX children are Autonomous-only.
Why these changes existed
The 0.7.13 Autonomous path could finish and still be wrong, or never finish.
The production incident that drove 0.7.14 was an office extraction: Autonomous mode, a 65K shared window, nine declared invoices, 19 tools, and response_format set. That shape is common. What happened was not:
| What you saw | What was actually going on |
|---|---|
| The agent "kept working" | Compaction shrank only the copy sent to the LLM. The live transcript stayed fat, so emergency compaction re-fired every turn toward the 300-call default. |
[recall_error] forever |
Compaction dropped the recall refs, then the masker told the model to recall_tool_result(ref=...) with nothing left to recall. |
| Six of nine invoices missing | The Critic was fed a 12K package that only named three files. It FAILed any answer with more records. The Refiner deleted the "unsupported" six. The generator had been right. |
ResultStatus.ERROR on a valid JSON reply |
The provider parsed two JSON objects as a schema failure before the harness could repair. A coverage follow-up child also died in preflight (19 tools but max 8 calls). |
| Children "forgot" the parent's 65K window | Each child was a blank Standard agent on the 8192 last-resort floor, plus a hidden 15-call plugin. |
Task.context did nothing |
The field was accepted and silently dropped. Callers pasted facts into objective, where they also reached the classifier and the Critic twice. |
Those are framework bugs, not prompt bugs. 0.7.14 fixes the harness so the same job terminates with completed, a schema-valid payload, and 9/9 resources — including against a Critic that literally counts headers in its own prompt.
What improved for you
| Before 0.7.14 | After 0.7.14 |
|---|---|
A run could loop until max_tool_calls |
Named stop: no_progress, context_tool_budget, deadline, … |
| You guessed why it stopped | result.termination_reason + result.diagnostics.explain() |
response_format was a hint |
Contract: result.parsed / result.schema_valid. Prose gets one tools-free finalizer. |
| Critic could overwrite a correct answer | I-10: FAIL on a partial view is UNCERTAIN. COMPLEX Critic sees the synthesizer's hand-off. |
Documents listed in objective were hope |
Task.resources is classified, tracked, and followed up once. |
| Children re-read everything on 8K | Children inherit the parent window, remaining time, and a layered evidence store. |
| Timeouts existed on paper | max_execution_time is enforced. llm_call_timeout / step_timeout only if you set them. |
| A 400 from an over-long prompt ended the run | Compact, write back, shrink the reply budget, then fail with context_overflow and numbers. |
result.output / str(result) still return the raw text. Existing Agent(...) code keeps working. New fields are additive.
How a run works
preflight ──► classify ──► SIMPLE loop | COMPLEX children
│
▼
coverage follow-up (at most once)
│
▼
schema finalizer (if response_format)
│
▼
Critic ──► Refiner (if quality check)
│
▼
AgentResult (termination_reason + diagnostics + parsed)
- Preflight —
working = window − reserve − system − tool schemas. Below 16K working tokens Autonomous downgrades to Standard (preflight_downgrade=True). - Classify — SIMPLE vs COMPLEX.
enable_decomposition=Falseskips this LLM call and still runs Critic/Refiner.Task.resourcesgrounds Gate 4: same sources / one record → SIMPLE. - Execute — SIMPLE is one tool loop. COMPLEX children inherit window, remaining wall clock, parent task fields, and a layered evidence view; findings merge back.
- Cover — unprocessed
Task.resourcesget exactly one bounded follow-up. - Finalize schema — when
response_formatis set, prose synthesis is skipped; one tools-free schema call runs instead. - Critic / Refiner — independent check. Partial evidence cannot FAIL an answer on its own.
- Return — always
termination_reasonanddiagnostics. Typed payload isresult.parsed.
Example — grounded extraction
This is the job the harness is built for: a known file list, tools that read them, a typed deliverable.
import asyncio
from pydantic import BaseModel
from nucleusiq.agents import Agent
from nucleusiq.agents.config import AgentConfig, ExecutionMode
from nucleusiq.agents.task import Task
from nucleusiq.prompts.zero_shot import ZeroShotPrompt
from nucleusiq.tools.decorators import tool
from nucleusiq_openai import BaseOpenAI
class InvoiceRecord(BaseModel):
vendor: str
total: float
currency: str
class InvoiceBatch(BaseModel):
invoices: list[InvoiceRecord]
@tool
def read_invoice(path: str) -> str:
"""Read one invoice file and return its text."""
catalog = {
"invoices/acme.pdf": "Vendor: ACME GmbH. Total: 1200.00 EUR.",
"invoices/globex.pdf": "Vendor: Globex. Total: 450.50 EUR.",
}
return catalog.get(path, f"missing: {path}")
async def main():
agent = Agent(
name="invoice-reader",
prompt=ZeroShotPrompt().configure(
system="Extract every declared invoice. Use tools. Return only the schema.",
),
llm=BaseOpenAI(model_name="gpt-4.1-mini"),
tools=[read_invoice],
response_format=InvoiceBatch,
config=AgentConfig(
execution_mode=ExecutionMode.AUTONOMOUS,
require_quality_check=True,
max_tool_calls=20,
coverage_followup=True,
),
)
result = await agent.execute(
Task(
id="inv-1",
objective="Extract vendor and total from every invoice.",
resources=["invoices/acme.pdf", "invoices/globex.pdf"],
context={"currency": "EUR"},
)
)
print(result.termination_reason) # completed
print(result.schema_valid) # True
print(result.parsed) # InvoiceBatch(invoices=[...])
print(result.output) # raw JSON — str(result) unchanged
print(result.diagnostics.explain())
print(result.diagnostics.coverage)
asyncio.run(main())
Do not treat result as the schema
execute() always returns AgentResult. The Pydantic instance is result.parsed. result.title is wrong.
Task.context is rendered as a bounded ## Task Context / ## Resources block. Do not paste the same list into objective — the classifier, Critic, and validator would see it twice.
Example — inspect a run
diagnostics is always on. You do not need enable_tracing for the report (tracing still fills llm_calls / tool_calls).
result = await agent.execute(task)
print(result.termination_reason)
print(result.diagnostics.explain())
for finding in result.diagnostics.findings:
print(finding.code, finding.title)
print(" cause:", finding.cause)
print(" now: ", finding.fix_now)
if result.diagnostics.coverage:
print(result.diagnostics.coverage)
for child in result.diagnostics.children:
print(child)
# Safe to paste into an issue — no task text, hashed ids
safe = result.diagnostics.redacted()
Offline:
python -m nucleusiq.agents.diagnostics report.json
python -m nucleusiq.agents.diagnostics report.json --json --redact
Exit 2 = critical, 1 = warning, 0 = clean. Finding catalog: v0.7.14 release notes.
Example — do not split a single record
If every child would re-read the same files, or the deliverable is one JSON object, a COMPLEX split loses coverage. Skip the classifier and keep Critic/Refiner:
config = AgentConfig(
execution_mode=ExecutionMode.AUTONOMOUS,
require_quality_check=True,
enable_decomposition=False, # no classifier LLM call
)
max_sub_agents=1 used to have the same routing effect but still paid for the classifier. Gate 4 now downgrades that split to SIMPLE automatically when Task.resources is set.
Example — refuse an incomplete extract
Default coverage records a gap. To withhold the answer after the one follow-up:
from nucleusiq.agents.task import Task
config = AgentConfig(
execution_mode=ExecutionMode.AUTONOMOUS,
coverage_followup=True,
evidence_gate_enforce=True, # residual gap → ABSTAINED
)
result = await agent.execute(
Task(
id="inv-all",
objective="Extract every invoice.",
resources=["invoices/acme.pdf", "invoices/globex.pdf", "invoices/initech.pdf"],
)
)
if result.is_abstained:
print(result.abstention_code) # coverage_incomplete
print(result.abstention_reason) # names the unprocessed files
print(result.output) # best candidate, kept for inspection
STANDARD records coverage and does not retry. SIMPLE Autonomous gets one retry prompt naming only the missing files (never on the last attempt). COMPLEX spawns one coverage-followup child on the unprocessed slice.
Every change — why it was needed, what improved
1. Named stop + always-on report
Why. A hung or "successful" run with a hollow answer was undiagnosable. Tracing was optional and too heavy to leave on.
What improved. Every exit sets result.termination_reason. result.diagnostics (RunReport, also agent.last_run_report) records resolved window/caps, decisions, counters, children, coverage, and analyzer findings — no task text. Independent of enable_tracing.
| Reason | You hit this when |
|---|---|
completed |
Accepted answer |
tool_budget |
Business max_tool_calls exhausted |
context_tool_budget |
Recall / workspace / corpus cap exhausted |
no_progress |
Identical tool rounds |
deadline |
max_execution_time exceeded |
llm_timeout |
You set llm_call_timeout and it expired |
context_overflow |
Provider 400 after recovery |
schema_invalid |
Final text still failed response_format |
critic_abstain |
Critic withheld the answer |
context_overflow, deadline, and llm_timeout are sticky — a later generic error does not overwrite them. Full catalog: release notes.
2. The tool loop cannot spin
Why. Compaction used to shrink only the request copy. The next turn re-compacted the same fat transcript. Recall refs vanished. The first user message was pinned, so the resource list and objective were the first things dropped. Unseen tool results were evicted; the model invented vendors.
What improved.
- No-progress guard — two identical rounds (or only duplicate /
[recall_error]banners) nudge once; a third stops withno_progress. Polling tools whose results change never trip it. max_context_tool_calls— recall / workspace / evidence / corpus stay exempt frommax_tool_callsbut now have their own cap (default2 × max_tool_calls).- Write-back — the reduced
prepare()view becomes the live transcript. - Recall catalog —
[evidence available for recall]keeps refs andargs={"path": "..."}. - Task head pinned — every leading system/user message before the first assistant turn stays.
- Adaptive offload — oversized current-turn results become recallable previews instead of silent eviction (
EVIDENCE_EVICTED_UNSEENif something still drops).
3. response_format is a contract
Why. The loop asked for JSON and then ran a prose synthesis pass ("write the full deliverable"). Short JSON failed the word threshold and came back as markdown. Provider-side parse errors (Extra data: line 2) aborted the run before repair.
What improved.
- Every content exit is validated (
StructuredOutputContract). - Prose / partial JSON → one tools-free finalizer (constrained decoding can actually fire).
- Synthesis is skipped when a schema is set.
- Autonomous adds a
schemavalidation layer between L1 and L2; retry prompts include the exact validator errors. - Critic: JSON-for-schema is the deliverable; prose is FAIL. Refiner returns only corrected JSON.
- Sub-agents never inherit
response_format. - Provider parse failure is retried once without provider parsing (
structured_parse_fallback). - Failed schema keeps
status=SUCCESS(you decide) and setstermination_reason="schema_invalid".
print(result.parsed) # InvoiceBatch(...) or None
print(result.schema_valid) # True / False / None
print(result.structured) # {schema, mode, valid, finalizer_runs, errors}
4. Window-derived budgets and preflight
Why. Caps were hardcoded for 128K ([:2000] findings, 12K Critic package). On 65K, children handed over the first 2 000 characters. On 8K, the 8 192 reply reserve exceeded the window (utilization 100 %, emergency compaction every turn). Children ignored the parent's ContextConfig and sized against 8192. Critic/Refiner caps used llm.get_context_window() (Ollama's 8192) even when you set 65K.
What improved.
BudgetResolversizes each hand-off from the resolved window. 128K keeps today's ceilings; 65K children send full findings; 8K shrinks.response_reservestays 8 192 when it fits; on 8K it becomes ~2 560.- Construction rejects an explicit reserve ≥ window (defaults never rejected).
- Undeclared window: one warning,
window_is_fallback=truein the report. - Preflight: Autonomous < 16K working tokens → Standard; < 32K →
marginal. Standard warns below 8K. - Critic/Refiner caps use
ContextConfig→ engine → provider → fallback.
Set an explicit window so children inherit it:
from nucleusiq.agents.context import ContextConfig, ContextStrategy
config = AgentConfig(
execution_mode=ExecutionMode.AUTONOMOUS,
context=ContextConfig(
optimal_budget=40_000,
strategy=ContextStrategy.PROGRESSIVE,
),
)
Override a child window only with sub_agent_context.
5. Time and overflow
Why. max_execution_time was documented and unused. llm_call_timeout=90 would have killed reasoning models if it had been real. A provider 400 for an over-long prompt ended the run. A stream exception tore down the caller's generator.
What improved.
max_execution_time(default 3600,0= unlimited) is enforced. Stop withdeadline, then synthesis / finalizer so you keep the best answer. Children inherit remaining seconds. SimpleRunner will not start a new Critic attempt after the deadline — it returns the best candidate.llm_call_timeout/step_timeoutfire only when you set them. Defaults are documentation.LLMTimeoutError→llm_timeout. A timed-out tool becomes an error the model can read; the loop continues.- Over-long prompt: emergency compact + write-back, then shrink
max_output_tokens, thencontext_overflowwith numbers — not a bare 400. execute_streamturns runtime errors into a terminalERRORevent and a report.
6. Task.resources and Gate 4
Why. "Process every invoice" in the objective was invisible to the classifier. COMPLEX splits spawned children that each re-read the same nine files, or that owned no files at all. Task.context was dropped.
What improved.
- Typed
Task.resources(context["resources"]is fallback).task.effective_resources()normalises. - Context / resources are rendered once, bounded, in front of the objective.
- Gate 4: same sources or one record → SIMPLE. Each COMPLEX sub-task must declare
"resources": [...]. - Coverage contract: every resource owned, no resource-less child, at most
decomposition_max_owners_per_resource(default 1) owners. Violation → SIMPLE (decomposition_downgraded). - Touch tracking matches tool args and result heads (
a.pdf↔docs/a.pdf).
7. COMPLEX children that share evidence
Why. Children started empty, on 8K, with a hidden 15-call cap. They saw only {id, objective}. Three children re-read nine files; the parent saw none of it. A follow-up budgeted at 2 × unprocessed died because 19 inherited tools exceeded that call cap. Failed children rendered as "".
What improved.
| Before | After |
|---|---|
| Blank Standard + 8192 | Inherit parent window / remaining time |
ModelCallLimitPlugin(15) |
Gone. Inherit max_tool_calls (unset → 80) |
{id, objective} |
Full child Task (attachments, context, resource slice) |
| Empty stores | Layered parent view; merge local writes back |
| Silent child death | error on the finding and in diagnostics.children |
| Follow-up preflight fail | _aux_tool_budget never below the tool list |
decomposition_gather_first=True (opt-in, idempotent=True tools) fetches every resource once before analysis children start.
8. Honest Critic (I-10)
Why. The Critic is an independent model. If you show it less evidence than the generator used, it will FAIL a correct answer. gpt-oss 120B counted three invoice headers in its prompt and deleted the other six. Gemma 4 COMPLEX verified nine records against an empty package (pass 1.00 + CRITIC_PARTIAL_VIEW) because synthesis made no tool calls of its own.
What improved. Invariant: the verifier sees at least what the generator saw, or knows it is seeing less.
- Synthesis package sized from
BudgetResolver(65K Critic package 12K → 40K). - Omission drops whole items and says the list is incomplete — never a mid-list cut that looks like a shorter complete list.
- Required Resources Processed section (harness-verified, from tool traffic).
- Critic gets package and rehydrated raw trace. Partial view →
## EVIDENCE VISIBILITY. FAIL → UNCERTAIN (critic_fail_downgraded). PASS is never touched. - Refiner may not drop "unsupported" records unless the tool results contradict them.
- COMPLEX: the exact findings text the synthesizer read is shown as
## MATERIAL THE GENERATOR WAS GIVEN. A tool-less generator whose inputs are all shown is a complete view.
9. WebSearchTool
Why. Autonomous research jobs needed a local search tool that works on every provider, not only hosted web_search.
What improved. from nucleusiq.tools import WebSearchTool — DuckDuckGo by default, no API key. Title / URL / snippet only. See Web search.
Reading AgentResult
result = await agent.execute(task)
result.output # raw text
str(result) # same
result.status # SUCCESS | ERROR | HALTED | ABSTAINED
result.termination_reason # why it stopped
result.parsed # schema instance (or None)
result.schema_valid # True / False / None
result.diagnostics # RunReport
result.diagnostics.coverage # if Task.resources was set
result.metadata["coverage"] # same numbers
result.autonomous # Critic verdicts when tracing is on
Config that matters
from nucleusiq.agents.config import AgentConfig, ExecutionMode
config = AgentConfig(
execution_mode=ExecutionMode.AUTONOMOUS,
require_quality_check=True,
enable_decomposition=True,
decomposition_max_owners_per_resource=1,
decomposition_gather_first=False,
coverage_followup=True,
evidence_gate_enforce=False,
max_tool_calls=80,
max_context_tool_calls=None, # 2 × max_tool_calls
max_execution_time=3600, # always on; 0 = unlimited
# llm_call_timeout=120, # uncomment to enforce
# step_timeout=90,
preflight_downgrade=True,
preflight_min_working_tokens=16_000,
sub_agent_context=None,
)
Field-by-field: Agent config.
Common mistakes
| Mistake | What to do instead |
|---|---|
print(result.title) after response_format=Summary |
print(result.parsed.title) |
Paste the file list into objective |
Task.resources=[...] |
Expect Task.context to stay private |
It is sent to the model (bounded) |
Assume default llm_call_timeout=90 is a timer |
It is not, until you set the field |
| Split one JSON record across children | enable_decomposition=False or let Gate 4 stay SIMPLE |
| Give children a smaller window by accident | Do not set sub_agent_context unless you mean it |
Treat ABSTAINED as ERROR |
Check result.is_abstained / abstention_code |
| Copy Ollama dict-args onto OpenAI | OpenAI / Groq still want a JSON string |
Copy Anthropic additionalProperties: false onto Gemini |
Gemini strips that keyword on purpose |
See also
- Autonomous workflow — copy-paste invoice example
- Execution modes — Direct vs Standard vs Autonomous
- Tasks —
resourcesand renderedcontext - Structured output —
parsed/schema_valid - Observability —
RunReport/ analyzer CLI - v0.7.14 release notes — every harness item
- Production path — when Autonomous is worth the cost