Example: Autonomous Workflow
Copy-paste recipe for Autonomous mode after the v0.7.14 harness work.
Why the APIs look like this (Critic parity, coverage, named stops, result.parsed): Autonomous mode guide. This page is the runnable example.
What the harness now guarantees
These are framework rules, not prompt tips:
| Guarantee | What you see |
|---|---|
| Named stop | result.termination_reason — completed, tool_budget, no_progress, deadline, schema_invalid, critic_abstain, … |
| Always-on report | result.diagnostics (RunReport) — counters, decisions, children, coverage, findings. No task text. Independent of enable_tracing. |
| Enforced schema | Agent(response_format=Schema) is parsed and validated. Use result.parsed / result.schema_valid. result.output stays the raw text. |
| Verifier parity | The Critic sees the generator's inputs, or an honest incomplete-view notice. A FAIL on a partial view is downgraded to UNCERTAIN. |
| Coverage | Declare Task.resources. One bounded follow-up on gaps. Optional abstain via evidence_gate_enforce=True. |
llm_call_timeout / step_timeout are enforced only when you set them explicitly. max_execution_time (default 3600 s) is always enforced.
Invoice extraction (copy-paste)
This is the shape the harness is built for: a known set of documents, tools that read them, and a typed deliverable.
import asyncio
from pydantic import BaseModel
from nucleusiq.agents import Agent
from nucleusiq.agents.config import AgentConfig, ExecutionMode
from nucleusiq.agents.task import Task
from nucleusiq.prompts.zero_shot import ZeroShotPrompt
from nucleusiq.tools.decorators import tool
from nucleusiq_openai import BaseOpenAI
class InvoiceRecord(BaseModel):
vendor: str
total: float
currency: str
class InvoiceBatch(BaseModel):
invoices: list[InvoiceRecord]
@tool
def read_invoice(path: str) -> str:
"""Read one invoice file and return its text."""
catalog = {
"invoices/acme.pdf": "Vendor: ACME GmbH. Total: 1200.00 EUR.",
"invoices/globex.pdf": "Vendor: Globex. Total: 450.50 EUR.",
}
return catalog.get(path, f"missing: {path}")
async def main():
agent = Agent(
name="invoice-reader",
prompt=ZeroShotPrompt().configure(
system="Extract every declared invoice. Use tools. Return only the schema.",
),
llm=BaseOpenAI(model_name="gpt-4.1-mini"),
tools=[read_invoice],
response_format=InvoiceBatch,
config=AgentConfig(
execution_mode=ExecutionMode.AUTONOMOUS,
require_quality_check=True,
max_tool_calls=20,
enable_decomposition=True,
coverage_followup=True,
),
)
result = await agent.execute(
Task(
id="inv-1",
objective="Extract vendor and total from every invoice.",
resources=["invoices/acme.pdf", "invoices/globex.pdf"],
context={"currency": "EUR"},
)
)
print(result.termination_reason) # completed
print(result.schema_valid) # True
print(result.parsed) # InvoiceBatch(...)
print(result.output) # raw JSON text — str(result) unchanged
report = result.diagnostics
print(report.explain()) # markdown for an issue
for finding in report.findings:
print(finding.code, finding.title)
if report.coverage:
print(report.coverage)
asyncio.run(main())
Task.context is now actually sent to the model (a bounded ## Task Context / ## Resources block). Do not paste the same facts into objective unless you want the classifier, Critic, and validator to see them twice.
execute() always returns an AgentResult. Never treat result itself as the Pydantic schema — that is result.parsed.
What happens under the hood
- Preflight — Working tokens = window − reserve − system − tool schemas. Unfit Autonomous windows downgrade to Standard (
preflight_downgrade=True). - Decomposition — The classifier decides SIMPLE vs COMPLEX.
enable_decomposition=Falseskips that call and still runs Critic/Refiner.Task.resourcesgrounds Gate 4 (same sources / one record → SIMPLE). - Execution — SIMPLE runs one tool loop. COMPLEX children inherit the parent's window and remaining wall clock, share a layered evidence view, and merge findings back.
- Coverage — Unprocessed resources get exactly one bounded follow-up (default on).
- Critic — Independent verification. Partial evidence cannot FAIL an answer on its own (downgraded to UNCERTAIN).
- Refiner — Corrects gaps. Told not to drop records the Critic called unsupported unless the tool results contradict them.
- Structured finalizer — When
response_formatis set, prose synthesis is skipped; one tools-free schema call runs instead. - Return —
termination_reason+diagnosticsare always set.
COMPLEX children
When the classifier splits the job, children used to start as a blank Standard agent on an 8192 floor with a hidden 15-call plugin. In 0.7.14 they:
- inherit the parent's window (or
sub_agent_context) and remaining wall clock - inherit
max_tool_calls/max_retries(unset parent → Standard 80, not Autonomous 300) - receive a
Taskwith parent attachments, context, and theirresourcesslice - read through a layered parent evidence view; local writes merge back
- keep
erroron the finding if they fail (diagnostics.children[].error)
decomposition_gather_first=True (opt-in, needs idempotent=True tools) fetches every resource once before analysis children start.
Why the Critic cannot delete a correct answer
Invariant I-10: the verifier sees at least what the generator saw, or knows it is seeing less.
- The Critic receives the synthesis package and the raw tool trace (or, on COMPLEX, the exact sub-agent hand-off the synthesizer read).
- Partial views prepend
## EVIDENCE VISIBILITY. - A FAIL on a partial view is downgraded to UNCERTAIN — it can request another pass but cannot, on its own, force abstention or strip records.
- The Refiner is told not to remove records the Critic called unsupported unless the tool results contradict them.
Live gpt-oss 120B used to drop six of nine invoices because it only saw three headers in its prompt. That path now returns 9/9.
Inspect the run
print(result.termination_reason)
print(result.diagnostics.explain())
for child in result.diagnostics.children:
print(child)
if result.diagnostics.coverage:
print(result.diagnostics.coverage)
# Offline, no task text:
# python -m nucleusiq.agents.diagnostics report.json --redact
result.diagnostics.redacted() hashes identifiers so a report is safe to paste into an issue.
Config knobs that matter
from nucleusiq.agents.config import AgentConfig, ExecutionMode
config = AgentConfig(
execution_mode=ExecutionMode.AUTONOMOUS,
enable_decomposition=True, # False = skip classifier; still run Critic/Refiner
decomposition_max_owners_per_resource=1,
decomposition_gather_first=False, # opt-in gather child (needs idempotent tools)
coverage_followup=True, # one bounded retry on unprocessed Task.resources
evidence_gate_enforce=False, # True = ABSTAINED / coverage_incomplete
max_context_tool_calls=None, # None = 2 × max_tool_calls
max_execution_time=3600, # always enforced; 0 = unlimited
llm_call_timeout=90, # enforced only when set explicitly
step_timeout=60,
preflight_downgrade=True, # unfit Autonomous window → STANDARD
sub_agent_context=None, # None = inherit parent window
)
result.diagnostics is always populated. enable_tracing still controls the detailed llm_calls / tool_calls trace.
With Gemini
The same agent code works with Gemini — swap the LLM:
from nucleusiq_gemini import BaseGemini
agent = Agent(
name="invoice-reader",
prompt=ZeroShotPrompt().configure(
system="Extract every declared invoice. Use tools. Return only the schema.",
),
llm=BaseGemini(model_name="gemini-2.5-pro"),
tools=[read_invoice],
response_format=InvoiceBatch,
config=AgentConfig(
execution_mode=ExecutionMode.AUTONOMOUS,
require_quality_check=True,
),
)
With context management
For long-running Autonomous tasks, set an explicit window so Critic/Refiner hand-offs and children inherit it:
from nucleusiq.agents.context import ContextConfig, ContextStrategy
config = AgentConfig(
execution_mode=ExecutionMode.AUTONOMOUS,
require_quality_check=True,
context=ContextConfig(
optimal_budget=40_000,
strategy=ContextStrategy.PROGRESSIVE,
),
)
Hand-off caps (findings, Critic evidence, Refiner summary) now come from BudgetResolver — they scale with the resolved window instead of hardcoded 2000 / 8000 / 12000.
See also
- Autonomous mode guide — why each change existed and what improved
- Execution modes — Direct vs Standard vs Autonomous + harness table
- Tasks —
resourcesand renderedcontext - Structured output —
result.parsed/result.schema_valid - Agent config — new knobs
- v0.7.14 release notes
- Gemini quickstart — Autonomous mode with Gemini
- Context management — Context window management guide
- Production path — When to use Autonomous mode