Skip to content

Execution modes

NucleusIQ uses the Gearbox Strategy — three execution modes that scale from simple chat to autonomous reasoning. All modes work identically with any LLM provider (OpenAI, Gemini, Anthropic, Groq, Ollama, OpenAI-compatible).

v0.7.14 hardens the shared Standard / Autonomous tool loop so a run always stops for a named reason, cannot loop on compaction or recall, enforces response_format, and gives the Critic the same evidence the generator saw (or an honest "this view is incomplete" notice).

Comparison

Capability Direct Standard Autonomous
Memory Yes Yes Yes
Plugins Yes Yes Yes
Tools Yes (default 25 / run) Yes (default 80 / run) Yes (default 300 / run)
Tool loop Yes Yes Yes
Streaming Yes Yes Yes
Structured output Yes Yes Yes
Context management Yes Yes Yes
Synthesis pass No Yes Yes
Task decomposition No No Yes
Critic (verification) No No Yes
Refiner (correction) No No Yes
Validation pipeline No No Yes

Default tool-call budgets (and the maximum number of user-registered tools on the agent, excluding framework recall tools) are 25 / 80 / 300 for Direct / Standard / Autonomous. Override any time with AgentConfig(max_tool_calls=N) or read the resolved value via config.get_effective_max_tool_calls(). Context-management tools (recall_*, workspace, evidence, corpus) have a separate cap: max_context_tool_calls (default 2 × max_tool_calls).

Harness guarantees (v0.7.14)

Standard and Autonomous share one tool loop. These guarantees apply to both unless noted.

Guarantee What you see
Named stop result.termination_reason — completed, tool_budget, no_progress, deadline, schema_invalid, critic_abstain, context_overflow, …
Always-on report result.diagnostics (RunReport) — counters, decisions, children, coverage, findings. No task text. Independent of enable_tracing.
No infinite loop Identical tool rounds trip a no-progress guard. Context tools have their own cap. Compaction writes the reduced transcript back.
Enforced schema Agent(response_format=Schema) is parsed and validated. Prose triggers one tools-free finalizer. Use result.parsed / result.schema_valid.
Window-derived budgets Hand-off caps (findings, Critic evidence, Refiner summary) come from BudgetResolver, not hardcoded 2000 / 8000 / 12000.
Verifier parity (I-10) Autonomous only. The Critic sees the generator's inputs, or an honest incomplete-view notice. A FAIL on a partial view is downgraded to UNCERTAIN.
Coverage Autonomous. Task.resources is rendered, classified, and reconciled. One bounded follow-up on gaps. Optional abstain via evidence_gate_enforce=True.
print(result.termination_reason)
print(result.diagnostics.explain())

llm_call_timeout / step_timeout are enforced only when you set them explicitly. max_execution_time (default 3600 s) is always enforced.

Direct

Fast, single LLM call. Best for Q&A, classification, and simple lookups.

import asyncio
from nucleusiq.agents import Agent
from nucleusiq.agents.config import AgentConfig, ExecutionMode
from nucleusiq.prompts.zero_shot import ZeroShotPrompt
from nucleusiq_openai import BaseOpenAI

async def main():
    agent = Agent(
        name="classifier",
        prompt=ZeroShotPrompt().configure(
            system="Classify the user's message as: positive, negative, or neutral. Respond with only the label.",
        ),
        llm=BaseOpenAI(model_name="gpt-4.1-mini"),
        config=AgentConfig(execution_mode=ExecutionMode.DIRECT),
    )
    result = await agent.execute({"id": "d1", "objective": "I love this product, it's amazing!"})
    print(result.output)  # "positive"

asyncio.run(main())

Use Direct when latency matters and the task doesn't need tools or iteration.

Standard (default)

Tool-enabled, linear execution. The agent calls tools iteratively until it has enough information to answer.

import asyncio
from nucleusiq.agents import Agent
from nucleusiq.agents.config import AgentConfig, ExecutionMode
from nucleusiq.prompts.zero_shot import ZeroShotPrompt
from nucleusiq.tools.decorators import tool
from nucleusiq_openai import BaseOpenAI

@tool
def get_stock_price(symbol: str) -> str:
    """Get current stock price for a ticker symbol."""
    prices = {"AAPL": "$195.50", "GOOGL": "$175.20", "MSFT": "$420.30"}
    return prices.get(symbol.upper(), f"Unknown: {symbol}")

@tool
def calculate(expression: str) -> str:
    """Evaluate a math expression."""
    return str(eval(expression))

async def main():
    agent = Agent(
        name="finance-bot",
        prompt=ZeroShotPrompt().configure(
            system="You are a financial assistant. Use tools to look up stock prices and do calculations.",
        ),
        llm=BaseOpenAI(model_name="gpt-4.1-mini"),
        tools=[get_stock_price, calculate],
        config=AgentConfig(execution_mode=ExecutionMode.STANDARD),
    )
    result = await agent.execute({
        "id": "s1",
        "objective": "What is the total value of 100 shares of AAPL and 50 shares of GOOGL?",
    })
    print(result.output)
    print(f"Duration: {result.duration_ms}ms")

asyncio.run(main())

Use Standard for most production workflows: data lookup, calculations, file operations, API calls.

Synthesis pass

New in v0.7.6

After multiple rounds of tool calls, Standard mode makes one final LLM call without tools and with an explicit instruction to write the full deliverable. This breaks the "mode inertia" pattern where the model stays in tool-calling behaviour and returns a terse summary instead of full output.

How it works:

  1. Agent enters tool loop — calls tools, receives results, decides next action
  2. When the LLM stops calling tools, the framework detects the loop has ended
  3. If enable_synthesis=True (default), one final LLM call is made without tools
  4. The final call includes an explicit instruction to produce the complete deliverable
  5. The LLM writes the full response without being distracted by tool availability
config = AgentConfig(
    execution_mode=ExecutionMode.STANDARD,
    enable_synthesis=True,  # Default: True
)

Set enable_synthesis=False only if your agent's task is simple enough that synthesis adds unnecessary latency.

Autonomous

Planning, multi-step execution with Critic/Refiner verification. The agent decomposes complex tasks, executes subtasks, and verifies its own output.

import asyncio
from nucleusiq.agents import Agent
from nucleusiq.agents.config import AgentConfig, ExecutionMode
from nucleusiq.prompts.zero_shot import ZeroShotPrompt
from nucleusiq.tools.decorators import tool
from nucleusiq_openai import BaseOpenAI

@tool
def analyze_risk(factor: str) -> str:
    """Analyze a specific risk factor."""
    risks = {
        "market": "HIGH — 30% revenue impact in downturn",
        "supply_chain": "MEDIUM — 2-3 month disruption tolerance",
        "regulatory": "LOW — compliant with current frameworks",
    }
    return risks.get(factor, f"No data for: {factor}")

async def main():
    agent = Agent(
        name="risk-analyst",
        prompt=ZeroShotPrompt().configure(
            system=(
                "You are a risk analyst. Break down complex assessments into subtasks. "
                "Analyze each risk factor using tools, then provide a comprehensive report "
                "with severity ratings and mitigation recommendations."
            ),
        ),
        llm=BaseOpenAI(model_name="gpt-4.1-mini"),
        tools=[analyze_risk],
        config=AgentConfig(
            execution_mode=ExecutionMode.AUTONOMOUS,
            require_quality_check=True,
            max_iterations=5,
            max_tool_calls=50,
        ),
    )
    result = await agent.execute({
        "id": "a1",
        "objective": "Assess market, supply chain, and regulatory risks for Q4.",
    })
    print(result.output)
    print(result.termination_reason)
    print(result.diagnostics.explain())

    # Autonomous detail (when tracing is enabled)
    if result.autonomous:
        print(f"Sub-tasks: {result.autonomous.sub_task_names}")

asyncio.run(main())

Use Autonomous for complex, high-stakes tasks where correctness matters more than speed. Declare Task.resources when the job is about a known set of documents — the Decomposer, coverage follow-up, and Critic all use that list.

Canonical write-up (why 0.7.14 existed, what improved, examples): Autonomous mode. Copy-paste: Autonomous workflow.

What happens under the hood

  1. Preflight — Working tokens = window − reserve − system − tool schemas. Unfit Autonomous windows downgrade to Standard (preflight_downgrade=True).
  2. Decomposition — The classifier decides SIMPLE vs COMPLEX. enable_decomposition=False skips that call and still runs Critic/Refiner. Task.resources grounds Gate 4 (same sources / one record → SIMPLE).
  3. Execution — SIMPLE runs one tool loop. COMPLEX children inherit the parent's window and remaining wall clock, share a layered evidence view, and merge findings back.
  4. Coverage — Unprocessed Task.resources get exactly one bounded follow-up (default on).
  5. Critic — Independent verification. Partial evidence cannot FAIL an answer on its own (downgraded to UNCERTAIN).
  6. Refiner — Corrects gaps. Told not to drop records the Critic called unsupported unless the tool results contradict them.
  7. Structured finalizer — When response_format is set, prose synthesis is skipped; one tools-free schema call runs instead.
  8. Return — result.termination_reason + result.diagnostics are always set. result.output / str(result) stay the raw text.

COMPLEX children (v0.7.14)

Children no longer start as a blank Standard agent on an 8192 floor.

  • Inherit the parent's ContextConfig (or already-resolved window) and remaining wall clock. Set sub_agent_context only to override.
  • Inherit max_tool_calls and max_retries. The hardcoded ModelCallLimitPlugin(max_calls=15) is gone. An unset parent gives the child the Standard default 80, not Autonomous 300.
  • The child Task carries parent attachments, context (minus resources), metadata, and its resources slice.
  • Reads fall through a layered parent evidence view; local writes merge back after each child.
  • decomposition_gather_first=True (opt-in) fetches every resource once before analysis children start (needs idempotent=True tools).
  • Gather / coverage-follow-up children get a budget that is never below the tool list they were handed, so they cannot die in tool-count preflight.
  • Failed children keep their error text (diagnostics.children[].error).

Inspect the tree on result.diagnostics.children and result.diagnostics.explain().

Stop reasons and the analyzer

Every run sets result.termination_reason. Common values: completed, tool_budget, context_tool_budget, no_progress, deadline, llm_timeout, context_overflow, schema_invalid, critic_abstain. Full catalog: v0.7.14 release notes.

print(result.diagnostics.explain())
# or offline:
# python -m nucleusiq.agents.diagnostics report.json

Sub-agent telemetry rollup

New in v0.7.6

In Autonomous mode, sub-agent LLM calls, tool calls, and context telemetry are merged into the parent agent's AgentResult. v0.7.14 also records per-child window, reserve, max_tool_calls, resources, and errors on result.diagnostics.children.

Context management across modes

New in v0.7.6

All three modes support context window management via ContextConfig. Add it to prevent context overflow in tool-heavy agents:

from nucleusiq.agents.context import ContextConfig, ContextStrategy

config = AgentConfig(
    execution_mode=ExecutionMode.STANDARD,
    context=ContextConfig(
        optimal_budget=50_000,
        strategy=ContextStrategy.PROGRESSIVE,
    ),
)

See Context management for details.

Choosing a mode

Scenario Recommended mode Why
Simple Q&A, classification DIRECT Fastest, single LLM call
Tool-enabled data workflows STANDARD Iterative tool use, synthesis pass
Complex analysis with verification AUTONOMOUS Decomposition + Critic/Refiner
Chat applications STANDARD Good balance of capability and speed
Code generation with validation AUTONOMOUS Critic verifies correctness
Latency-sensitive endpoints DIRECT Minimal overhead
Research tasks with many sources STANDARD + context mgmt Handles large context

See also