v0.7.13 — OpenAI-compatible provider
Why v0.7.13 matters
The community ask was self-hosted and bring-your-own-model. v0.7.13 adds nucleusiq-openai-compatible: one adapter for every Chat Completions server. Core now declares provider identity instead of guessing it from the class name. All 32 Dependabot advisories are cleared.
Packages
| Package | Version | Notes |
|---|---|---|
nucleusiq |
0.7.13 | BaseLLM.PROVIDER_NAME; dead supports_native_output() removed |
nucleusiq-openai-compatible |
0.1.0 | New, Stable. BYOM / BYOK. Requires nucleusiq>=0.7.13 |
nucleusiq-openai |
0.7.1 | Responses API usage: input_tokens / output_tokens → prompt_tokens / completion_tokens |
nucleusiq-mcp |
0.1.1 | Security floor mcp>=1.28.1 |
nucleusiq-gemini |
0.3.1 | PROVIDER_NAME + declared pydantic |
nucleusiq-anthropic |
0.2.1 | PROVIDER_NAME + declared httpx / pydantic |
nucleusiq-ollama |
0.2.1 | same |
nucleusiq-groq |
0.1.1 | same |
Only the new provider floors on nucleusiq>=0.7.13. The others stay on >=0.7.12 — PROVIDER_NAME is inert on older cores.
4,382 tests passing across the monorepo.
What users can do now
pip install "nucleusiq>=0.7.13" nucleusiq-openai-compatible
from nucleusiq_openai_compatible import OpenAICompatibleLLM
llm = OpenAICompatibleLLM(
base_url="http://gpu-node-1:8000/v1",
model="gemma-4-27b-it",
context_window=32_768,
engine="vllm",
)
Drop that llm into any existing Agent. Tools, modes, memory, plugins, and the context engine do not change.
→ OpenAI-compatible provider · Quickstart
Core — declared provider identity
get_provider_from_llm() reads BaseLLM.PROVIDER_NAME first. Class-name matching is only a fallback for older adapters.
Without the declaration, OpenAICompatibleLLM would match as "openai" (the name contains that substring) and send OpenAI-cloud structured output to a self-hosted server.
supports_native_output() is gone. It hardcoded stale OpenAI model prefixes and was dead: both AUTO branches returned NATIVE. It was never a public export. OutputMode.AUTO still resolves to NATIVE. NATIVE means "hand the schema to the adapter", not "the server implements json_schema". Degradation is the adapter's job.
OpenAI 0.7.1 — Responses usage
On the Responses API path only (hosted tools, reasoning models):
- Non-streaming
call()no longer reports 0 tokens / $0.00 - Streaming maps
input_tokens/output_tokenstoprompt_tokens/completion_tokens
Live-endpoint fixes (new provider)
Found against a real Chat Completions server; mocks could not catch them:
- Flat tool-calls from core are nested under
functionbefore send (second tool-loop request no longer 400s). - Streaming COMPLETE now includes
tool_calls(the tool loop actually runs). - Inbound
response_formatalways goes through the policy layer (vLLM tools + schema trap).
The Ollama /v1 preset sets supports_json_schema=False. The shim accepts a schema and ignores it. Use nucleusiq-ollama for native format.
Security and packaging
nucleusiq-mcprequiresmcp>=1.28.1(session auth, task isolation, WebSocket Host/Origin).- Every provider now declares
httpx/pydanticit actually imports.openaiis capped at<3.0until the Responses path is live-verified. - Gemini README examples no longer use invalid
Agent(model=..., instructions=...)(community report #38).
Upgrade
pip install -U "nucleusiq>=0.7.13"
pip install nucleusiq-openai-compatible # only if you need self-hosted / BYOM
Existing OpenAI / Gemini / Anthropic / Groq / Ollama / MCP apps keep working on nucleusiq>=0.7.12. Add 0.7.13 when you adopt the new provider.
No public API break in core. PROVIDER_NAME is additive.
See also
- Changelog
- Monorepo CHANGELOG.md
- GitHub release: v0.7.13