Skip to content

v0.7.13 — OpenAI-compatible provider

Why v0.7.13 matters

The community ask was self-hosted and bring-your-own-model. v0.7.13 adds nucleusiq-openai-compatible: one adapter for every Chat Completions server. Core now declares provider identity instead of guessing it from the class name. All 32 Dependabot advisories are cleared.

Packages

Package Version Notes
nucleusiq 0.7.13 BaseLLM.PROVIDER_NAME; dead supports_native_output() removed
nucleusiq-openai-compatible 0.1.0 New, Stable. BYOM / BYOK. Requires nucleusiq>=0.7.13
nucleusiq-openai 0.7.1 Responses API usage: input_tokens / output_tokens → prompt_tokens / completion_tokens
nucleusiq-mcp 0.1.1 Security floor mcp>=1.28.1
nucleusiq-gemini 0.3.1 PROVIDER_NAME + declared pydantic
nucleusiq-anthropic 0.2.1 PROVIDER_NAME + declared httpx / pydantic
nucleusiq-ollama 0.2.1 same
nucleusiq-groq 0.1.1 same

Only the new provider floors on nucleusiq>=0.7.13. The others stay on >=0.7.12 — PROVIDER_NAME is inert on older cores.

4,382 tests passing across the monorepo.

What users can do now

pip install "nucleusiq>=0.7.13" nucleusiq-openai-compatible
from nucleusiq_openai_compatible import OpenAICompatibleLLM

llm = OpenAICompatibleLLM(
    base_url="http://gpu-node-1:8000/v1",
    model="gemma-4-27b-it",
    context_window=32_768,
    engine="vllm",
)

Drop that llm into any existing Agent. Tools, modes, memory, plugins, and the context engine do not change.

→ OpenAI-compatible provider · Quickstart

Core — declared provider identity

get_provider_from_llm() reads BaseLLM.PROVIDER_NAME first. Class-name matching is only a fallback for older adapters.

Without the declaration, OpenAICompatibleLLM would match as "openai" (the name contains that substring) and send OpenAI-cloud structured output to a self-hosted server.

supports_native_output() is gone. It hardcoded stale OpenAI model prefixes and was dead: both AUTO branches returned NATIVE. It was never a public export. OutputMode.AUTO still resolves to NATIVE. NATIVE means "hand the schema to the adapter", not "the server implements json_schema". Degradation is the adapter's job.

OpenAI 0.7.1 — Responses usage

On the Responses API path only (hosted tools, reasoning models):

  • Non-streaming call() no longer reports 0 tokens / $0.00
  • Streaming maps input_tokens / output_tokens to prompt_tokens / completion_tokens

Live-endpoint fixes (new provider)

Found against a real Chat Completions server; mocks could not catch them:

  1. Flat tool-calls from core are nested under function before send (second tool-loop request no longer 400s).
  2. Streaming COMPLETE now includes tool_calls (the tool loop actually runs).
  3. Inbound response_format always goes through the policy layer (vLLM tools + schema trap).

The Ollama /v1 preset sets supports_json_schema=False. The shim accepts a schema and ignores it. Use nucleusiq-ollama for native format.

Security and packaging

  • nucleusiq-mcp requires mcp>=1.28.1 (session auth, task isolation, WebSocket Host/Origin).
  • Every provider now declares httpx / pydantic it actually imports. openai is capped at <3.0 until the Responses path is live-verified.
  • Gemini README examples no longer use invalid Agent(model=..., instructions=...) (community report #38).

Upgrade

pip install -U "nucleusiq>=0.7.13"
pip install nucleusiq-openai-compatible   # only if you need self-hosted / BYOM

Existing OpenAI / Gemini / Anthropic / Groq / Ollama / MCP apps keep working on nucleusiq>=0.7.12. Add 0.7.13 when you adopt the new provider.

No public API break in core. PROVIDER_NAME is additive.

See also