v0.7.14 — Autonomous harness hardening
Why v0.7.14 matters
Autonomous mode is now a bounded, diagnosable harness. Every run ends for a named reason, the tool loop cannot spin on compaction or recall, response_format is enforced (not just requested), and the Critic sees the same evidence the generator saw — or is told it is seeing less. Core also ships WebSearchTool. Two provider patches fix bugs only those SDKs hit.
These are framework guarantees, not prompt tips. They apply in Standard and Autonomous (they share the tool loop) unless a section says Autonomous-only.
Packages
| Package | Version | Notes |
|---|---|---|
nucleusiq |
0.7.14 | Harness hardening + WebSearchTool |
nucleusiq-ollama |
0.2.2 | Multi-turn tool arguments sent as a mapping (official ollama SDK) |
nucleusiq-anthropic |
0.2.2 | Nested structured-output schemas set additionalProperties: false on every object |
nucleusiq-openai |
0.7.1 | Unchanged |
nucleusiq-openai-compatible |
0.1.0 | Unchanged |
nucleusiq-gemini |
0.3.1 | Unchanged |
nucleusiq-groq |
0.1.1 | Unchanged |
nucleusiq-mcp |
0.1.1 | Unchanged |
Provider floors stay as they were (nucleusiq>=0.7.12, openai-compatible >=0.7.13).
pip install -U "nucleusiq>=0.7.14" nucleusiq-ollama==0.2.2 nucleusiq-anthropic==0.2.2
What strengthened the Agent harness
| Guarantee | What you see |
|---|---|
| Named stop | result.termination_reason — see Stop reasons |
| Always-on report | result.diagnostics (RunReport) — counters, decisions, children, coverage, analyzer findings. No task text. Independent of enable_tracing. |
| No infinite loop | No-progress guard, separate context-tool cap, write-back compaction, recall catalog, task-head pin |
| Enforced schema | Agent(response_format=Schema) is parsed and validated. Prose triggers one tools-free finalizer. Use result.parsed / result.schema_valid. |
| Window-derived budgets | Hand-off caps come from BudgetResolver, not hardcoded 2000 / 8000 / 12000 |
| Verifier parity (I-10) | The Critic sees the generator's inputs, or an honest incomplete-view notice. A FAIL on a partial view is downgraded to UNCERTAIN. |
| Coverage | Task.resources is rendered, classified, and reconciled. One bounded follow-up on gaps. Optional abstain with evidence_gate_enforce=True. |
Example — Autonomous run you can inspect
import asyncio
from pydantic import BaseModel
from nucleusiq.agents import Agent
from nucleusiq.agents.config import AgentConfig, ExecutionMode
from nucleusiq.agents.task import Task
from nucleusiq.prompts.zero_shot import ZeroShotPrompt
from nucleusiq.tools.decorators import tool
from nucleusiq_openai import BaseOpenAI
class InvoiceRecord(BaseModel):
vendor: str
total: float
currency: str
class InvoiceBatch(BaseModel):
invoices: list[InvoiceRecord]
@tool
def read_invoice(path: str) -> str:
"""Read one invoice file and return its text."""
catalog = {
"invoices/acme.pdf": "Vendor: ACME GmbH. Total: 1200.00 EUR.",
"invoices/globex.pdf": "Vendor: Globex. Total: 450.50 EUR.",
}
return catalog.get(path, f"missing: {path}")
async def main():
agent = Agent(
name="invoice-reader",
prompt=ZeroShotPrompt().configure(
system="Extract every declared invoice. Use tools. Return only the schema.",
),
llm=BaseOpenAI(model_name="gpt-4.1-mini"),
tools=[read_invoice],
response_format=InvoiceBatch,
config=AgentConfig(
execution_mode=ExecutionMode.AUTONOMOUS,
require_quality_check=True,
max_tool_calls=20,
enable_decomposition=True,
coverage_followup=True,
),
)
result = await agent.execute(
Task(
id="inv-1",
objective="Extract vendor and total from every invoice.",
resources=["invoices/acme.pdf", "invoices/globex.pdf"],
context={"currency": "EUR"},
)
)
print(result.termination_reason) # completed
print(result.schema_valid) # True
print(result.parsed) # InvoiceBatch(...)
print(result.output) # raw JSON text — str(result) unchanged
report = result.diagnostics
print(report.explain()) # markdown for an issue
for finding in report.findings:
print(finding.code, finding.title)
if report.coverage:
print(report.coverage)
asyncio.run(main())
Task.context is now actually sent to the model (a bounded ## Task Context / ## Resources block). Do not paste the same facts into objective unless you want the classifier, Critic, and validator to see them twice.
execute() always returns an AgentResult. Never treat result itself as the Pydantic schema — that is result.parsed.
Stop reasons
Every exit of the Standard / streaming / Autonomous tool loops, the Autonomous dispatcher, and Agent.execute() sets result.termination_reason. The last exit wins, so a retried-then-successful run reads completed.
| Reason | Meaning |
|---|---|
completed |
Finished with an accepted answer |
structured_output |
Stopped because the schema contract was satisfied |
tool_budget |
Business max_tool_calls exhausted |
context_tool_budget |
Recall / workspace / evidence / corpus cap exhausted |
no_progress |
Identical tool rounds (or only duplicate / [recall_error] banners) |
emergency_compaction |
Circuit-breaker after three emergency events in one run |
deadline |
max_execution_time exceeded |
llm_timeout |
Explicit llm_call_timeout expired |
context_overflow |
Provider 400 for an over-long prompt after recovery |
critic_abstain |
Autonomous Critic withheld the answer |
schema_invalid |
Final text still failed response_format |
plugin_halt |
A plugin raised PluginHalt |
empty_response |
Model returned nothing usable |
refusal |
Model refused the request |
error |
Unhandled exception |
context_overflow, deadline, and llm_timeout are sticky: the generic error recorded by execute() never overwrites them.
Run report and analyzer
result.diagnostics is a RunReport (also agent.last_run_report). Always on. Summary-level. No task text, tool arguments, or payloads. Independent of enable_tracing.
What it carries:
- Resolved config — window, reserve, working budget, caps, tool counts,
window_is_fallback - Decisions — classification, preflight, gather-first, coverage, critic views, synthesis hand-off
- Counters — rounds, LLM calls, business vs context-management tools, dedup banners, recall errors, stalled rounds, compactions, emergency count, write-backs, unseen evictions
- Children — per-child window, reserve,
max_tool_calls, resources, touches, errors - Coverage —
resources,touched,unprocessed,acknowledged,unaccounted,complete,followup,enforce,blocked - Termination record + a bounded timeline
- Analyzer findings — sorted critical → warning → info
print(result.diagnostics.explain()) # markdown for an issue
print(result.diagnostics.redacted()) # hashes identifiers; safe to share
print(result.diagnostics.analyze()) # same findings, offline
Offline CLI (accepts a report JSON, - for stdin, or a whole AgentResult.summary()):
python -m nucleusiq.agents.diagnostics report.json
python -m nucleusiq.agents.diagnostics report.json --json --redact
Exit status: 2 critical / 1 warning / 0 clean.
Each finding has code, title, evidence (the exact report fields used), cause, fix_now, fix_release. AgentResult.display() shows the top three.
| Code | Typical cause |
|---|---|
LOOP_RECALL_DEADLOCK |
Recall errors / missing refs |
LOOP_NO_PROGRESS |
Identical tool rounds |
EMERGENCY_THRASH |
Repeated emergency compaction |
WINDOW_FALLBACK |
Context window assumed 128K |
CHILD_WINDOW_MISMATCH / CHILD_HIDDEN_CAP |
Child sized smaller than the parent |
DECOMP_SHARED_SOURCE / DECOMP_DOWNGRADED |
Gate 4 / coverage contract |
DECOMP_COVERAGE_GAP / COVERAGE_GAP |
Declared resources not processed |
HANDOFF_TRUNCATED |
Child findings cut for space |
CHILDREN_FAILED |
A child errored (error text is kept) |
PREFLIGHT_UNFIT / PREFLIGHT_MARGINAL |
Working-token budget |
RESERVE_LT_MAX_TOKENS |
Reply reserve smaller than llm_max_output_tokens |
CONTEXT_OVERFLOW_400 |
Provider rejected an over-long prompt |
TOOL_BUDGET_EXHAUSTED / CONTEXT_TOOL_BUDGET_EXHAUSTED |
Caps |
DEADLINE / LLM_TIMEOUT |
Time |
CRITIC_ABSTAIN / CRITIC_PARTIAL_VIEW |
Verifier withheld, or saw less than the generator |
SYNTH_VS_STRUCTURED / SCHEMA_NOT_SATISFIED |
Schema contract |
EMPTY_RESPONSES / GATHER_SKIPPED |
Empty model / gather-first skipped |
EVIDENCE_EVICTED_UNSEEN |
Tool results dropped before the model read them |
Public API: analyze(report), attach_findings(report), recommendations(findings), load_report(...), registered_rules(), RunReport.analyze().
The tool loop cannot spin
No-progress guard
Each round is signed (tool names + arguments + result hashes). Two identical rounds, or two rounds made only of duplicate-call banners / [recall_error], inject a one-time nudge: answer from existing evidence. A third stops with no_progress. Polling tools whose results change never trip it. A guard stop still gets the tools-free synthesis (or structured finalizer) chance.
Separate context-tool budget
recall_*, workspace, evidence, and corpus tools stay exempt from max_tool_calls (that quota is for your external actions) but now have max_context_tool_calls (default 2 × max_tool_calls). Exhaustion stops with context_tool_budget.
Write-back compaction
prepare() used to shrink only the copy sent to the LLM. The live transcript stayed fat, so the next turn re-triggered the same emergency compaction (it could loop toward Autonomous's 300-call default). Any reduced view is now written back as the continuing transcript. Three emergency events in one run still trip a last-resort circuit breaker (emergency_compaction).
Recall stays addressable
Conversation compaction (Tier 2) and emergency reduction (Tier 3) used to drop store-backed receipts with the rest of the old groups. Dedup then told the model to recall_tool_result(ref=...) with no ref left — [recall_error] forever. Both tiers now emit one [evidence available for recall] catalog ([observation consumed] and [context_ref:] receipts). Each catalog line carries the call's arguments (args={"path": "..."}) so the model can tell which input a ref belongs to.
Task head is pinned
Both compaction tiers used to pin only the first user message. Prompt templates emit several leading user messages (preamble, Task.resources, objective), so under pressure the model "forgot" what it still had to process. The whole task head — leading system messages plus every user message before the first assistant turn — is now pinned.
Unseen evidence is not evicted silently
A round whose tool results exceed the working budget before any assistant turn has read them used to go straight to emergency eviction (live Gemma-4 then invented vendors and totals). Tier-1 now adaptively offloads raw tool results to the store as recallable previews with an explicit recall_tool_result(ref=...) hint. What still had to be dropped is counted (unseen_evidence_evicted) and reported as EVIDENCE_EVICTED_UNSEEN.
[context_ref:] receipts are rehydrated for the Critic, Refiner, and synthesis pass — not just [observation consumed] markers.
Enforced structured output
Agent(response_format=Schema) is now a contract, not a request.
- Every content exit of the Standard / streaming / Autonomous tool loops is parsed and validated (
StructuredOutputContract). - Prose or partial JSON triggers one tools-free structured finalizer (no
toolson the request, so constrained decoding can enforce the schema server-side). - A loop stopped by a budget or guard runs the finalizer instead of the prose synthesis pass.
- The prose synthesis pass is skipped whenever a schema is set — JSON is almost always under the word threshold, and the old "write the full deliverable" nudge replaced it with markdown.
- Autonomous adds a deterministic
schemavalidation layer between L1 and L2 and feeds the validator's exact errors into the retry prompt. The Critic prompt has an Output Contract section (JSON-for-schema is the deliverable; prose is FAIL). The Refiner is told to return only the corrected JSON. - Sub-agents never inherit
response_format.
Read the result this way:
| Field | Meaning |
|---|---|
result.output / str(result) |
Raw text (unchanged). Also mirrored at metadata["raw_output"]. |
result.parsed |
Validated schema instance, or None |
result.schema_valid |
True / False when a schema was set, else None |
result.structured |
{"schema", "mode", "valid", "finalizer_runs", "errors"} |
A final answer that still fails the schema keeps status=SUCCESS (you decide) but sets termination_reason="schema_invalid". Runs without response_format see no change (parsed=None, structured=None).
Provider-side parse fallback. In NATIVE mode the adapter parses inside llm.call. Extra JSON, JSON + prose, or a field that fails validation used to raise StructuredOutputError before the contract saw the text, and the run ended in ERROR. The harness now catches that once, strips the schema type while keeping the wire format, retries, and lets the contract decide. Timeline event: structured_parse_fallback.
Window-derived budgets
Hand-off caps used to be hardcoded for 128K ([:2000] findings, 8 000 / 4 000 Refiner, 12 000 Critic package). BudgetResolver (ContextEngine.budgets) now sizes each role from the resolved window (overhead + share of prompt budget + floor + ceiling).
- 128K: today's caps stay the upper bound (findings ≤ 40K chars, Refiner candidate ≤ 32K).
- 65K: three children hand over their whole findings instead of the first 2 000 chars.
- 8K: caps shrink instead of overflowing.
RunReport.config_resolved["budgets"] shows the numbers.
Derived response_reserve. An explicit value always wins. The 8 192 default is kept when it fits (≤ ¼ window and ≥ max_output_tokens + 512), so 128K users see no change. On an 8K model the old fixed 8 192 exceeded the window (utilization pinned at 100 %, emergency compaction every turn); it now resolves to 2 560. The report shows config_resolved["response_reserve"].
Construction-time validation. AgentConfig rejects an explicit ContextConfig.response_reserve >= max_context_tokens, and an explicit llm_max_output_tokens > response_reserve. Defaults are never rejected.
Undeclared window. When neither ContextConfig.max_context_tokens nor the provider gives a real number, the agent warns once (assuming 128000) and the report sets config_resolved["window_is_fallback"] = true.
Critic / Refiner caps use the resolved window (ContextConfig → engine → provider → fallback), not a flat provider default such as Ollama's 8192. A user-configured 65K window no longer starves the Critic to the 500-char floor.
Preflight fitness
Before the first LLM call the agent measures system-prompt and tool-schema tokens:
working = window − response_reserve − system − tool_schemas
| Mode | Threshold | Default | Effect |
|---|---|---|---|
| Autonomous | preflight_min_working_tokens |
16 000 | Downgrade to Standard (result.mode == "standard", decisions["preflight"]["action"] == "downgraded_to_standard"). Set preflight_downgrade=False to force Autonomous. |
| Autonomous | preflight_marginal_working_tokens |
32 000 | Run proceeds, flagged marginal |
| Standard | preflight_standard_min_working_tokens |
8 000 | Warning only |
The full breakdown lands in decisions["preflight"] and config_resolved["preflight"].
Time and overflow
Wall clock. max_execution_time (default 3600 s, 0 = unlimited) is finally enforced. The tool loops check a monotonic deadline at every boundary and stop with deadline, then run tools-free synthesis or the structured finalizer so you get the best answer gathered so far. Autonomous SimpleRunner refuses to start a new Critic/Refiner attempt once the budget is spent and returns the best candidate instead of abstaining. COMPLEX children inherit the parent's remaining seconds.
Opt-in per-call timeouts. llm_call_timeout (documented 90) and step_timeout (documented 60) had defaults for years without being enforced. They are now enforced only when you set them explicitly. Defaults stay inert so reasoning models and long tools keep working. A slow LLM call raises LLMTimeoutError (llm_timeout). A slow tool becomes an Error: Tool 'x' timed out after Ns result the model can reason about; the loop continues. Children inherit explicit values.
ContextLengthError recovery. A provider 400 for an over-long prompt is no longer terminal on the first try. The harness forces emergency compaction and writes the reduced transcript back, then on a second rejection shrinks max_output_tokens so prompt + reply fits. Only then does the run fail — with context_overflow and the numbers (window, prompt≈, max_output_tokens) in the report instead of a bare 400.
execute_stream never leaks. Runtime exceptions from a mode's stream become a terminal ERROR event (and a run report) instead of tearing down the consumer's generator. Configuration errors still raise at setup.
Task.resources and the 4-gate classifier
Task.resources is the typed list of documents / files / records a task must cover. context["resources"] is a fallback; the typed field wins. task.effective_resources() normalises either.
Task.context was previously accepted and silently dropped. MessageBuilder now emits one bounded ## Task Context / ## Resources (N) user message immediately before the objective (window-derived cap, explicit [... task context truncated] marker, resources beyond 60 summarised as a count). Tasks without context / resources produce the same messages as before.
Autonomous-only, when resources are set:
- Gate 4 — separable sources, separate outputs. A split where every child would re-read the same sources, or where the deliverable is one record, is SIMPLE. The classifier also sees already-indexed corpus documents and business tool names.
- Coverage contract on COMPLEX. Every resource owned by at least one sub-task, no resource-less sub-task, at most
decomposition_max_owners_per_resource(default 1) owners. Matching is case-insensitive and accepts a trailing path component (a.pdf↔docs/a.pdf). Any violation downgrades to SIMPLE (decisions["classification"],decomposition_downgraded). enable_decomposition=Falseskips the classifier LLM call and still runs validation, Critic, and Refiner.max_sub_agents=1had that routing effect before but still paid the classifier call.
COMPLEX children
Children no longer start as a blank Standard agent on an 8192 floor.
| Before | After |
|---|---|
Blank AgentConfig(STANDARD) — parent's 65K window dropped |
Inherit parent's ContextConfig, or the already-resolved engine window. Set sub_agent_context only to override. |
Hardcoded ModelCallLimitPlugin(max_calls=15) |
Plugin gone. Children inherit max_tool_calls and max_retries. Unset parent → Standard default 80, not Autonomous 300. |
Child task was {"id", "objective"} |
Child Task carries parent attachments, context (minus resources), metadata, and its resources slice. No slice → full list. |
| Empty store / corpus / dossier | Layered views: reads fall through to the parent, writes stay local. After each child, local entries merge into the parent (evidence_merged). |
Finding was {id, objective, result} only |
Also status, termination_reason, touched_resources, refs, merged, and error (failures are no longer silent). |
The child's system prompt says when shared evidence is already indexed (use search_document_corpus … before calling any read or fetch tool). Per-child window, reserve, max_tool_calls, and counters land in diagnostics.children.
decomposition_gather_first=True (default False) — opt-in read-only gather child (idempotent tools only, max_tool_calls = min(parent, 2 × len(resources)), no synthesis, remaining wall clock) fetches every resource into the shared stores before analysis children start. Skipped with a recorded reason when no resources or no idempotent tools are present.
Auxiliary children (gather, coverage follow-up) get a budget that is never below the tool list they were handed (_aux_tool_budget). A follow-up budgeted at 2 × unprocessed used to fail preflight (has 19 tools but STANDARD mode allows max 8) and report an empty CHILDREN_FAILED.
Coverage follow-up
Every run with Task.resources ends with unprocessed = resources − touched (tool traffic ∪ children ∪ indexed corpus documents) in result.metadata["coverage"] and diagnostics.coverage.
Before the answer is accepted, the harness spends exactly one bounded follow-up (coverage_followup=True, default; only when business tools are present):
- COMPLEX — a
coverage-followupchild on exactly the unprocessed slice (max_tool_calls = min(parent, 2 × len(unprocessed)), shared evidence, appended as a finding so synthesis sees it). - SIMPLE Autonomous — one retry prompt naming only the missing resources. Never on the last attempt, so a gap cannot turn the only answer into nothing.
- STANDARD — records coverage without retrying.
evidence_gate_enforce=True then applies to resources: a residual gap returns status=ABSTAINED, abstention_code="coverage_incomplete", the unprocessed names in abstention_reason, and the output kept for inspection. Default stays False (record, don't block).
RunReport.redacted() hashes nested resource lists so a shared report cannot leak a path.
Verifier parity (I-10)
Design invariant: the verifier sees at least what the generator saw, or knows it is seeing less.
Live gpt-oss 120B counted invoice headers in its own prompt, FAILed any answer with more records, and the Refiner deleted six of nine correct rows. Four changes close that:
- Budget-derived synthesis package — sized per consumer (
synthesis_package,critic_evidence_total) instead of a fixed 12 000 chars. Workspace notes stay present (previews shrink first). 65K Critic package grows from 12K to 40K; 8K shrinks instead of overflowing. - Honest omission — whole items are dropped, never a list cut mid-item. The package appends
[N more … omitted for space — this list is INCOMPLETE; absence here is not absence of evidence]. A required Resources Processed section states, from tool traffic, how many declared resources were read and names them. - Critic gets the map and the territory — raw tool trace (rehydrated, capped from the resolved window minus the package) plus the package.
CriticView(package_complete,raw_trace_complete) lands indecisions["critic_views"]. A partial view prepends## EVIDENCE VISIBILITY. A FAIL on a partial view is downgraded to UNCERTAIN — the verifier can request another pass but cannot, on its own, condemn an answer or force abstention. Counters:critic_partial_views/critic_fail_downgraded. Finding:CRITIC_PARTIAL_VIEW. - Refiner honours the view — leads with the same Resources Processed facts. When the critique came from a partial view, it is forbidden to remove records the Critic called unsupported unless the tool results contradict them.
COMPLEX path. A synthesis over sub-agent findings usually makes no tool calls, so the Critic used to see an empty package and verify nine records against nothing (Gemma 4: pass 1.00 with CRITIC_PARTIAL_VIEW). The exact hand-off the synthesizer reads is now published as generator inputs (## MATERIAL THE GENERATOR WAS GIVEN). CriticView.complete means parity: the tool trace is shown whole (or there was none) or the package is a complete map. A tool-less generator whose inputs are all shown is a complete view.
Config knobs that matter
from nucleusiq.agents.config import AgentConfig, ExecutionMode
config = AgentConfig(
execution_mode=ExecutionMode.AUTONOMOUS,
enable_decomposition=True, # False = skip classifier; still run Critic/Refiner
decomposition_max_owners_per_resource=1,
decomposition_gather_first=False, # opt-in gather child (needs idempotent tools)
coverage_followup=True, # one bounded retry on unprocessed Task.resources
evidence_gate_enforce=False, # True = ABSTAINED / coverage_incomplete after a gap
max_context_tool_calls=None, # None = 2 × max_tool_calls
max_execution_time=3600, # always enforced; 0 = unlimited
llm_call_timeout=90, # enforced only when you set it explicitly
step_timeout=60, # same
preflight_downgrade=True, # unfit Autonomous window → STANDARD
preflight_min_working_tokens=16_000,
sub_agent_context=None, # None = inherit parent window
)
result.diagnostics is always populated. enable_tracing still controls the detailed llm_calls / tool_calls trace.
WebSearchTool
First-class core search (from nucleusiq.tools import WebSearchTool). Default backend is DuckDuckGo (no API key; ddgs is a core dependency). Results are title / URL / snippet only — no page fetch. Public http(s) only (blocks javascript:, file:, localhost, private IPs). Constructor-only allowed_domains / blocked_domains. Paid engines: google (API key + CSE id), brave, tavily, serper. bing uses the free Bing index through ddgs (Microsoft retired the official API). Custom adapters: WebSearchBackendFactory.register.
This is a local tool. Provider-hosted search (OpenAITool.web_search(), AnthropicTool.web_search(), GeminiTool.google_search()) is separate.
Provider patches (only these two)
Ollama 0.2.2 — the official SDK requires tool_calls[].function.arguments to be a dict. A JSON string failed validation on the second tool-loop call. OpenAI / Groq / openai-compatible still expect a string; do not copy this fix onto them.
Anthropic 0.2.2 — Claude's grammar requires additionalProperties: false on every object. Nested Pydantic models (list[Item]) used to 400. Gemini strips that keyword on purpose; do not apply the Anthropic closer there.
Upgrade
pip install -U "nucleusiq>=0.7.14"
Existing Agent(prompt=..., llm=..., config=...) code keeps working. New fields on AgentResult and Task are additive.
result.output/str(result)still return the raw text.result.parsedis set only whenresponse_formatis configured.Task.resourcesis optional. Without it, coverage reconciliation is a no-op.
See also
- Autonomous mode — why each change existed and what improved
- Execution modes — gearbox + harness loop
- Autonomous workflow — copy-paste example
- Tasks —
resourcesand renderedcontext - Structured output —
parsed/schema_valid - Agent config — new knobs
- Observability —
RunReport/ analyzer CLI - Web search —
WebSearchTool - Monorepo CHANGELOG.md — 0.7.14
- GitHub release: v0.7.14