Skip to content

Autonomous mode

Autonomous is the gearbox's third gear: decompose → execute → verify → correct. Use it when a wrong answer is worse than a slower one — multi-document extraction, high-stakes analysis, anything the Critic should be allowed to reject.

v0.7.14 turned that gearbox into a bounded harness. The model still reasons. The framework now decides when to stop, what the Critic is allowed to see, and whether every declared document was actually processed.

Copy-paste first: Autonomous workflow. This page is the full guide — why the changes existed, what improved, and how to use the new APIs.

When to use it

Use Autonomous when… Stay on Standard when…
The job is about a known set of documents One lookup + one answer
You need an independent Critic / Refiner Latency matters more than verification
A missed file is a wrong deliverable Coverage is "best effort"
You want a typed response_format that the loop must satisfy Free-form prose is the product

Standard and Autonomous share the same tool loop. Loop guards, termination_reason, diagnostics, and the structured-output contract apply to both. Decomposition, Critic/Refiner, coverage follow-up, and COMPLEX children are Autonomous-only.

Why these changes existed

The 0.7.13 Autonomous path could finish and still be wrong, or never finish.

The production incident that drove 0.7.14 was an office extraction: Autonomous mode, a 65K shared window, nine declared invoices, 19 tools, and response_format set. That shape is common. What happened was not:

What you saw What was actually going on
The agent "kept working" Compaction shrank only the copy sent to the LLM. The live transcript stayed fat, so emergency compaction re-fired every turn toward the 300-call default.
[recall_error] forever Compaction dropped the recall refs, then the masker told the model to recall_tool_result(ref=...) with nothing left to recall.
Six of nine invoices missing The Critic was fed a 12K package that only named three files. It FAILed any answer with more records. The Refiner deleted the "unsupported" six. The generator had been right.
ResultStatus.ERROR on a valid JSON reply The provider parsed two JSON objects as a schema failure before the harness could repair. A coverage follow-up child also died in preflight (19 tools but max 8 calls).
Children "forgot" the parent's 65K window Each child was a blank Standard agent on the 8192 last-resort floor, plus a hidden 15-call plugin.
Task.context did nothing The field was accepted and silently dropped. Callers pasted facts into objective, where they also reached the classifier and the Critic twice.

Those are framework bugs, not prompt bugs. 0.7.14 fixes the harness so the same job terminates with completed, a schema-valid payload, and 9/9 resources — including against a Critic that literally counts headers in its own prompt.

What improved for you

Before 0.7.14 After 0.7.14
A run could loop until max_tool_calls Named stop: no_progress, context_tool_budget, deadline, …
You guessed why it stopped result.termination_reason + result.diagnostics.explain()
response_format was a hint Contract: result.parsed / result.schema_valid. Prose gets one tools-free finalizer.
Critic could overwrite a correct answer I-10: FAIL on a partial view is UNCERTAIN. COMPLEX Critic sees the synthesizer's hand-off.
Documents listed in objective were hope Task.resources is classified, tracked, and followed up once.
Children re-read everything on 8K Children inherit the parent window, remaining time, and a layered evidence store.
Timeouts existed on paper max_execution_time is enforced. llm_call_timeout / step_timeout only if you set them.
A 400 from an over-long prompt ended the run Compact, write back, shrink the reply budget, then fail with context_overflow and numbers.

result.output / str(result) still return the raw text. Existing Agent(...) code keeps working. New fields are additive.

How a run works

preflight ──► classify ──► SIMPLE loop | COMPLEX children
                 │
                 ▼
         coverage follow-up (at most once)
                 │
                 ▼
         schema finalizer (if response_format)
                 │
                 ▼
         Critic ──► Refiner (if quality check)
                 │
                 ▼
         AgentResult (termination_reason + diagnostics + parsed)
  1. Preflight — working = window − reserve − system − tool schemas. Below 16K working tokens Autonomous downgrades to Standard (preflight_downgrade=True).
  2. Classify — SIMPLE vs COMPLEX. enable_decomposition=False skips this LLM call and still runs Critic/Refiner. Task.resources grounds Gate 4: same sources / one record → SIMPLE.
  3. Execute — SIMPLE is one tool loop. COMPLEX children inherit window, remaining wall clock, parent task fields, and a layered evidence view; findings merge back.
  4. Cover — unprocessed Task.resources get exactly one bounded follow-up.
  5. Finalize schema — when response_format is set, prose synthesis is skipped; one tools-free schema call runs instead.
  6. Critic / Refiner — independent check. Partial evidence cannot FAIL an answer on its own.
  7. Return — always termination_reason and diagnostics. Typed payload is result.parsed.

Example — grounded extraction

This is the job the harness is built for: a known file list, tools that read them, a typed deliverable.

import asyncio
from pydantic import BaseModel
from nucleusiq.agents import Agent
from nucleusiq.agents.config import AgentConfig, ExecutionMode
from nucleusiq.agents.task import Task
from nucleusiq.prompts.zero_shot import ZeroShotPrompt
from nucleusiq.tools.decorators import tool
from nucleusiq_openai import BaseOpenAI


class InvoiceRecord(BaseModel):
    vendor: str
    total: float
    currency: str


class InvoiceBatch(BaseModel):
    invoices: list[InvoiceRecord]


@tool
def read_invoice(path: str) -> str:
    """Read one invoice file and return its text."""
    catalog = {
        "invoices/acme.pdf": "Vendor: ACME GmbH. Total: 1200.00 EUR.",
        "invoices/globex.pdf": "Vendor: Globex. Total: 450.50 EUR.",
    }
    return catalog.get(path, f"missing: {path}")


async def main():
    agent = Agent(
        name="invoice-reader",
        prompt=ZeroShotPrompt().configure(
            system="Extract every declared invoice. Use tools. Return only the schema.",
        ),
        llm=BaseOpenAI(model_name="gpt-4.1-mini"),
        tools=[read_invoice],
        response_format=InvoiceBatch,
        config=AgentConfig(
            execution_mode=ExecutionMode.AUTONOMOUS,
            require_quality_check=True,
            max_tool_calls=20,
            coverage_followup=True,
        ),
    )
    result = await agent.execute(
        Task(
            id="inv-1",
            objective="Extract vendor and total from every invoice.",
            resources=["invoices/acme.pdf", "invoices/globex.pdf"],
            context={"currency": "EUR"},
        )
    )

    print(result.termination_reason)   # completed
    print(result.schema_valid)         # True
    print(result.parsed)               # InvoiceBatch(invoices=[...])
    print(result.output)               # raw JSON — str(result) unchanged
    print(result.diagnostics.explain())
    print(result.diagnostics.coverage)

asyncio.run(main())

Do not treat result as the schema

execute() always returns AgentResult. The Pydantic instance is result.parsed. result.title is wrong.

Task.context is rendered as a bounded ## Task Context / ## Resources block. Do not paste the same list into objective — the classifier, Critic, and validator would see it twice.

Example — inspect a run

diagnostics is always on. You do not need enable_tracing for the report (tracing still fills llm_calls / tool_calls).

result = await agent.execute(task)

print(result.termination_reason)
print(result.diagnostics.explain())

for finding in result.diagnostics.findings:
    print(finding.code, finding.title)
    print("  cause:", finding.cause)
    print("  now:  ", finding.fix_now)

if result.diagnostics.coverage:
    print(result.diagnostics.coverage)

for child in result.diagnostics.children:
    print(child)

# Safe to paste into an issue — no task text, hashed ids
safe = result.diagnostics.redacted()

Offline:

python -m nucleusiq.agents.diagnostics report.json
python -m nucleusiq.agents.diagnostics report.json --json --redact

Exit 2 = critical, 1 = warning, 0 = clean. Finding catalog: v0.7.14 release notes.

Example — do not split a single record

If every child would re-read the same files, or the deliverable is one JSON object, a COMPLEX split loses coverage. Skip the classifier and keep Critic/Refiner:

config = AgentConfig(
    execution_mode=ExecutionMode.AUTONOMOUS,
    require_quality_check=True,
    enable_decomposition=False,   # no classifier LLM call
)

max_sub_agents=1 used to have the same routing effect but still paid for the classifier. Gate 4 now downgrades that split to SIMPLE automatically when Task.resources is set.

Example — refuse an incomplete extract

Default coverage records a gap. To withhold the answer after the one follow-up:

from nucleusiq.agents.task import Task

config = AgentConfig(
    execution_mode=ExecutionMode.AUTONOMOUS,
    coverage_followup=True,
    evidence_gate_enforce=True,   # residual gap → ABSTAINED
)

result = await agent.execute(
    Task(
        id="inv-all",
        objective="Extract every invoice.",
        resources=["invoices/acme.pdf", "invoices/globex.pdf", "invoices/initech.pdf"],
    )
)

if result.is_abstained:
    print(result.abstention_code)     # coverage_incomplete
    print(result.abstention_reason)   # names the unprocessed files
    print(result.output)              # best candidate, kept for inspection

STANDARD records coverage and does not retry. SIMPLE Autonomous gets one retry prompt naming only the missing files (never on the last attempt). COMPLEX spawns one coverage-followup child on the unprocessed slice.

Every change — why it was needed, what improved

1. Named stop + always-on report

Why. A hung or "successful" run with a hollow answer was undiagnosable. Tracing was optional and too heavy to leave on.

What improved. Every exit sets result.termination_reason. result.diagnostics (RunReport, also agent.last_run_report) records resolved window/caps, decisions, counters, children, coverage, and analyzer findings — no task text. Independent of enable_tracing.

Reason You hit this when
completed Accepted answer
tool_budget Business max_tool_calls exhausted
context_tool_budget Recall / workspace / corpus cap exhausted
no_progress Identical tool rounds
deadline max_execution_time exceeded
llm_timeout You set llm_call_timeout and it expired
context_overflow Provider 400 after recovery
schema_invalid Final text still failed response_format
critic_abstain Critic withheld the answer

context_overflow, deadline, and llm_timeout are sticky — a later generic error does not overwrite them. Full catalog: release notes.

2. The tool loop cannot spin

Why. Compaction used to shrink only the request copy. The next turn re-compacted the same fat transcript. Recall refs vanished. The first user message was pinned, so the resource list and objective were the first things dropped. Unseen tool results were evicted; the model invented vendors.

What improved.

  • No-progress guard — two identical rounds (or only duplicate / [recall_error] banners) nudge once; a third stops with no_progress. Polling tools whose results change never trip it.
  • max_context_tool_calls — recall / workspace / evidence / corpus stay exempt from max_tool_calls but now have their own cap (default 2 × max_tool_calls).
  • Write-back — the reduced prepare() view becomes the live transcript.
  • Recall catalog — [evidence available for recall] keeps refs and args={"path": "..."}.
  • Task head pinned — every leading system/user message before the first assistant turn stays.
  • Adaptive offload — oversized current-turn results become recallable previews instead of silent eviction (EVIDENCE_EVICTED_UNSEEN if something still drops).

3. response_format is a contract

Why. The loop asked for JSON and then ran a prose synthesis pass ("write the full deliverable"). Short JSON failed the word threshold and came back as markdown. Provider-side parse errors (Extra data: line 2) aborted the run before repair.

What improved.

  • Every content exit is validated (StructuredOutputContract).
  • Prose / partial JSON → one tools-free finalizer (constrained decoding can actually fire).
  • Synthesis is skipped when a schema is set.
  • Autonomous adds a schema validation layer between L1 and L2; retry prompts include the exact validator errors.
  • Critic: JSON-for-schema is the deliverable; prose is FAIL. Refiner returns only corrected JSON.
  • Sub-agents never inherit response_format.
  • Provider parse failure is retried once without provider parsing (structured_parse_fallback).
  • Failed schema keeps status=SUCCESS (you decide) and sets termination_reason="schema_invalid".
print(result.parsed)            # InvoiceBatch(...) or None
print(result.schema_valid)      # True / False / None
print(result.structured)        # {schema, mode, valid, finalizer_runs, errors}

4. Window-derived budgets and preflight

Why. Caps were hardcoded for 128K ([:2000] findings, 12K Critic package). On 65K, children handed over the first 2 000 characters. On 8K, the 8 192 reply reserve exceeded the window (utilization 100 %, emergency compaction every turn). Children ignored the parent's ContextConfig and sized against 8192. Critic/Refiner caps used llm.get_context_window() (Ollama's 8192) even when you set 65K.

What improved.

  • BudgetResolver sizes each hand-off from the resolved window. 128K keeps today's ceilings; 65K children send full findings; 8K shrinks.
  • response_reserve stays 8 192 when it fits; on 8K it becomes ~2 560.
  • Construction rejects an explicit reserve ≥ window (defaults never rejected).
  • Undeclared window: one warning, window_is_fallback=true in the report.
  • Preflight: Autonomous < 16K working tokens → Standard; < 32K → marginal. Standard warns below 8K.
  • Critic/Refiner caps use ContextConfig → engine → provider → fallback.

Set an explicit window so children inherit it:

from nucleusiq.agents.context import ContextConfig, ContextStrategy

config = AgentConfig(
    execution_mode=ExecutionMode.AUTONOMOUS,
    context=ContextConfig(
        optimal_budget=40_000,
        strategy=ContextStrategy.PROGRESSIVE,
    ),
)

Override a child window only with sub_agent_context.

5. Time and overflow

Why. max_execution_time was documented and unused. llm_call_timeout=90 would have killed reasoning models if it had been real. A provider 400 for an over-long prompt ended the run. A stream exception tore down the caller's generator.

What improved.

  • max_execution_time (default 3600, 0 = unlimited) is enforced. Stop with deadline, then synthesis / finalizer so you keep the best answer. Children inherit remaining seconds. SimpleRunner will not start a new Critic attempt after the deadline — it returns the best candidate.
  • llm_call_timeout / step_timeout fire only when you set them. Defaults are documentation. LLMTimeoutError → llm_timeout. A timed-out tool becomes an error the model can read; the loop continues.
  • Over-long prompt: emergency compact + write-back, then shrink max_output_tokens, then context_overflow with numbers — not a bare 400.
  • execute_stream turns runtime errors into a terminal ERROR event and a report.

6. Task.resources and Gate 4

Why. "Process every invoice" in the objective was invisible to the classifier. COMPLEX splits spawned children that each re-read the same nine files, or that owned no files at all. Task.context was dropped.

What improved.

  • Typed Task.resources (context["resources"] is fallback). task.effective_resources() normalises.
  • Context / resources are rendered once, bounded, in front of the objective.
  • Gate 4: same sources or one record → SIMPLE. Each COMPLEX sub-task must declare "resources": [...].
  • Coverage contract: every resource owned, no resource-less child, at most decomposition_max_owners_per_resource (default 1) owners. Violation → SIMPLE (decomposition_downgraded).
  • Touch tracking matches tool args and result heads (a.pdf ↔ docs/a.pdf).

7. COMPLEX children that share evidence

Why. Children started empty, on 8K, with a hidden 15-call cap. They saw only {id, objective}. Three children re-read nine files; the parent saw none of it. A follow-up budgeted at 2 × unprocessed died because 19 inherited tools exceeded that call cap. Failed children rendered as "".

What improved.

Before After
Blank Standard + 8192 Inherit parent window / remaining time
ModelCallLimitPlugin(15) Gone. Inherit max_tool_calls (unset → 80)
{id, objective} Full child Task (attachments, context, resource slice)
Empty stores Layered parent view; merge local writes back
Silent child death error on the finding and in diagnostics.children
Follow-up preflight fail _aux_tool_budget never below the tool list

decomposition_gather_first=True (opt-in, idempotent=True tools) fetches every resource once before analysis children start.

8. Honest Critic (I-10)

Why. The Critic is an independent model. If you show it less evidence than the generator used, it will FAIL a correct answer. gpt-oss 120B counted three invoice headers in its prompt and deleted the other six. Gemma 4 COMPLEX verified nine records against an empty package (pass 1.00 + CRITIC_PARTIAL_VIEW) because synthesis made no tool calls of its own.

What improved. Invariant: the verifier sees at least what the generator saw, or knows it is seeing less.

  • Synthesis package sized from BudgetResolver (65K Critic package 12K → 40K).
  • Omission drops whole items and says the list is incomplete — never a mid-list cut that looks like a shorter complete list.
  • Required Resources Processed section (harness-verified, from tool traffic).
  • Critic gets package and rehydrated raw trace. Partial view → ## EVIDENCE VISIBILITY. FAIL → UNCERTAIN (critic_fail_downgraded). PASS is never touched.
  • Refiner may not drop "unsupported" records unless the tool results contradict them.
  • COMPLEX: the exact findings text the synthesizer read is shown as ## MATERIAL THE GENERATOR WAS GIVEN. A tool-less generator whose inputs are all shown is a complete view.

9. WebSearchTool

Why. Autonomous research jobs needed a local search tool that works on every provider, not only hosted web_search.

What improved. from nucleusiq.tools import WebSearchTool — DuckDuckGo by default, no API key. Title / URL / snippet only. See Web search.

Reading AgentResult

result = await agent.execute(task)

result.output                 # raw text
str(result)                   # same
result.status                 # SUCCESS | ERROR | HALTED | ABSTAINED
result.termination_reason     # why it stopped
result.parsed                 # schema instance (or None)
result.schema_valid           # True / False / None
result.diagnostics            # RunReport
result.diagnostics.coverage   # if Task.resources was set
result.metadata["coverage"]   # same numbers
result.autonomous             # Critic verdicts when tracing is on

Config that matters

from nucleusiq.agents.config import AgentConfig, ExecutionMode

config = AgentConfig(
    execution_mode=ExecutionMode.AUTONOMOUS,
    require_quality_check=True,
    enable_decomposition=True,
    decomposition_max_owners_per_resource=1,
    decomposition_gather_first=False,
    coverage_followup=True,
    evidence_gate_enforce=False,
    max_tool_calls=80,
    max_context_tool_calls=None,     # 2 × max_tool_calls
    max_execution_time=3600,         # always on; 0 = unlimited
    # llm_call_timeout=120,          # uncomment to enforce
    # step_timeout=90,
    preflight_downgrade=True,
    preflight_min_working_tokens=16_000,
    sub_agent_context=None,
)

Field-by-field: Agent config.

Common mistakes

Mistake What to do instead
print(result.title) after response_format=Summary print(result.parsed.title)
Paste the file list into objective Task.resources=[...]
Expect Task.context to stay private It is sent to the model (bounded)
Assume default llm_call_timeout=90 is a timer It is not, until you set the field
Split one JSON record across children enable_decomposition=False or let Gate 4 stay SIMPLE
Give children a smaller window by accident Do not set sub_agent_context unless you mean it
Treat ABSTAINED as ERROR Check result.is_abstained / abstention_code
Copy Ollama dict-args onto OpenAI OpenAI / Groq still want a JSON string
Copy Anthropic additionalProperties: false onto Gemini Gemini strips that keyword on purpose

See also