Skip to content

Example: Autonomous Workflow

Copy-paste recipe for Autonomous mode after the v0.7.14 harness work.

Why the APIs look like this (Critic parity, coverage, named stops, result.parsed): Autonomous mode guide. This page is the runnable example.

What the harness now guarantees

These are framework rules, not prompt tips:

Guarantee What you see
Named stop result.termination_reason — completed, tool_budget, no_progress, deadline, schema_invalid, critic_abstain, …
Always-on report result.diagnostics (RunReport) — counters, decisions, children, coverage, findings. No task text. Independent of enable_tracing.
Enforced schema Agent(response_format=Schema) is parsed and validated. Use result.parsed / result.schema_valid. result.output stays the raw text.
Verifier parity The Critic sees the generator's inputs, or an honest incomplete-view notice. A FAIL on a partial view is downgraded to UNCERTAIN.
Coverage Declare Task.resources. One bounded follow-up on gaps. Optional abstain via evidence_gate_enforce=True.

llm_call_timeout / step_timeout are enforced only when you set them explicitly. max_execution_time (default 3600 s) is always enforced.

Invoice extraction (copy-paste)

This is the shape the harness is built for: a known set of documents, tools that read them, and a typed deliverable.

import asyncio
from pydantic import BaseModel
from nucleusiq.agents import Agent
from nucleusiq.agents.config import AgentConfig, ExecutionMode
from nucleusiq.agents.task import Task
from nucleusiq.prompts.zero_shot import ZeroShotPrompt
from nucleusiq.tools.decorators import tool
from nucleusiq_openai import BaseOpenAI


class InvoiceRecord(BaseModel):
    vendor: str
    total: float
    currency: str


class InvoiceBatch(BaseModel):
    invoices: list[InvoiceRecord]


@tool
def read_invoice(path: str) -> str:
    """Read one invoice file and return its text."""
    catalog = {
        "invoices/acme.pdf": "Vendor: ACME GmbH. Total: 1200.00 EUR.",
        "invoices/globex.pdf": "Vendor: Globex. Total: 450.50 EUR.",
    }
    return catalog.get(path, f"missing: {path}")


async def main():
    agent = Agent(
        name="invoice-reader",
        prompt=ZeroShotPrompt().configure(
            system="Extract every declared invoice. Use tools. Return only the schema.",
        ),
        llm=BaseOpenAI(model_name="gpt-4.1-mini"),
        tools=[read_invoice],
        response_format=InvoiceBatch,
        config=AgentConfig(
            execution_mode=ExecutionMode.AUTONOMOUS,
            require_quality_check=True,
            max_tool_calls=20,
            enable_decomposition=True,
            coverage_followup=True,
        ),
    )
    result = await agent.execute(
        Task(
            id="inv-1",
            objective="Extract vendor and total from every invoice.",
            resources=["invoices/acme.pdf", "invoices/globex.pdf"],
            context={"currency": "EUR"},
        )
    )

    print(result.termination_reason)          # completed
    print(result.schema_valid)                # True
    print(result.parsed)                      # InvoiceBatch(...)
    print(result.output)                      # raw JSON text — str(result) unchanged

    report = result.diagnostics
    print(report.explain())                   # markdown for an issue
    for finding in report.findings:
        print(finding.code, finding.title)
    if report.coverage:
        print(report.coverage)

asyncio.run(main())

Task.context is now actually sent to the model (a bounded ## Task Context / ## Resources block). Do not paste the same facts into objective unless you want the classifier, Critic, and validator to see them twice.

execute() always returns an AgentResult. Never treat result itself as the Pydantic schema — that is result.parsed.

What happens under the hood

  1. Preflight — Working tokens = window − reserve − system − tool schemas. Unfit Autonomous windows downgrade to Standard (preflight_downgrade=True).
  2. Decomposition — The classifier decides SIMPLE vs COMPLEX. enable_decomposition=False skips that call and still runs Critic/Refiner. Task.resources grounds Gate 4 (same sources / one record → SIMPLE).
  3. Execution — SIMPLE runs one tool loop. COMPLEX children inherit the parent's window and remaining wall clock, share a layered evidence view, and merge findings back.
  4. Coverage — Unprocessed resources get exactly one bounded follow-up (default on).
  5. Critic — Independent verification. Partial evidence cannot FAIL an answer on its own (downgraded to UNCERTAIN).
  6. Refiner — Corrects gaps. Told not to drop records the Critic called unsupported unless the tool results contradict them.
  7. Structured finalizer — When response_format is set, prose synthesis is skipped; one tools-free schema call runs instead.
  8. Return — termination_reason + diagnostics are always set.

COMPLEX children

When the classifier splits the job, children used to start as a blank Standard agent on an 8192 floor with a hidden 15-call plugin. In 0.7.14 they:

  • inherit the parent's window (or sub_agent_context) and remaining wall clock
  • inherit max_tool_calls / max_retries (unset parent → Standard 80, not Autonomous 300)
  • receive a Task with parent attachments, context, and their resources slice
  • read through a layered parent evidence view; local writes merge back
  • keep error on the finding if they fail (diagnostics.children[].error)

decomposition_gather_first=True (opt-in, needs idempotent=True tools) fetches every resource once before analysis children start.

Why the Critic cannot delete a correct answer

Invariant I-10: the verifier sees at least what the generator saw, or knows it is seeing less.

  • The Critic receives the synthesis package and the raw tool trace (or, on COMPLEX, the exact sub-agent hand-off the synthesizer read).
  • Partial views prepend ## EVIDENCE VISIBILITY.
  • A FAIL on a partial view is downgraded to UNCERTAIN — it can request another pass but cannot, on its own, force abstention or strip records.
  • The Refiner is told not to remove records the Critic called unsupported unless the tool results contradict them.

Live gpt-oss 120B used to drop six of nine invoices because it only saw three headers in its prompt. That path now returns 9/9.

Inspect the run

print(result.termination_reason)
print(result.diagnostics.explain())
for child in result.diagnostics.children:
    print(child)
if result.diagnostics.coverage:
    print(result.diagnostics.coverage)

# Offline, no task text:
# python -m nucleusiq.agents.diagnostics report.json --redact

result.diagnostics.redacted() hashes identifiers so a report is safe to paste into an issue.

Config knobs that matter

from nucleusiq.agents.config import AgentConfig, ExecutionMode

config = AgentConfig(
    execution_mode=ExecutionMode.AUTONOMOUS,
    enable_decomposition=True,              # False = skip classifier; still run Critic/Refiner
    decomposition_max_owners_per_resource=1,
    decomposition_gather_first=False,       # opt-in gather child (needs idempotent tools)
    coverage_followup=True,                 # one bounded retry on unprocessed Task.resources
    evidence_gate_enforce=False,            # True = ABSTAINED / coverage_incomplete
    max_context_tool_calls=None,            # None = 2 × max_tool_calls
    max_execution_time=3600,                # always enforced; 0 = unlimited
    llm_call_timeout=90,                    # enforced only when set explicitly
    step_timeout=60,
    preflight_downgrade=True,               # unfit Autonomous window → STANDARD
    sub_agent_context=None,                 # None = inherit parent window
)

result.diagnostics is always populated. enable_tracing still controls the detailed llm_calls / tool_calls trace.

With Gemini

The same agent code works with Gemini — swap the LLM:

from nucleusiq_gemini import BaseGemini

agent = Agent(
    name="invoice-reader",
    prompt=ZeroShotPrompt().configure(
        system="Extract every declared invoice. Use tools. Return only the schema.",
    ),
    llm=BaseGemini(model_name="gemini-2.5-pro"),
    tools=[read_invoice],
    response_format=InvoiceBatch,
    config=AgentConfig(
        execution_mode=ExecutionMode.AUTONOMOUS,
        require_quality_check=True,
    ),
)

With context management

For long-running Autonomous tasks, set an explicit window so Critic/Refiner hand-offs and children inherit it:

from nucleusiq.agents.context import ContextConfig, ContextStrategy

config = AgentConfig(
    execution_mode=ExecutionMode.AUTONOMOUS,
    require_quality_check=True,
    context=ContextConfig(
        optimal_budget=40_000,
        strategy=ContextStrategy.PROGRESSIVE,
    ),
)

Hand-off caps (findings, Critic evidence, Refiner summary) now come from BudgetResolver — they scale with the resolved window instead of hardcoded 2000 / 8000 / 12000.

See also