Docs

Chat completions

Kunavo's /v1/chat/completions endpoint speaks OpenAI's chat completions protocol for the Claude and GPT families. Streaming, tools and JSON output work through the same SDK; what reasoning_effort does depends on the family, and this page says which.

Ta dokumentacja jest dostępna w języku angielskim. Polskiego przewodnika jeszcze nie mamy — o płatnościach przeczytasz tutaj:Płatności — ceny w USD, doładowanie salda BLIK-iem lub kartą →

Endpoint: POST /v1/chat/completions. Request and response shape match OpenAI's chat completions API exactly — including streaming and the optional tool_calls / reasoning_tokens fields.

Basic call

resp = client.chat.completions.create(
    model="claude-sonnet-5",
    messages=[
        {"role": "system", "content": "You are a senior staff engineer."},
        {"role": "user", "content": "Pros and cons of postgres LISTEN/NOTIFY for a job queue?"},
    ],
    max_tokens=800,
)
print(resp.choices[0].message.content)

Parameters

Every standard OpenAI parameter is accepted. Some only make sense for certain providers — the translator passes through what each upstream supports.

ParamTypeNotes
modelstring (required)Any enabled slug from /v1/models.
messagesarray (required)Standard OpenAI message format.
temperature0..2Sampling temperature. Default 1. Has no effect on the models listed under this table.
top_p0..1Nucleus sampling. Same exception as temperature.
max_tokensintOutput cap; max_completion_tokens wins when both are sent. On GPT models the reasoning tokens count against it. A Claude request that sends neither gets 4,096.
streamboolStream chunks as SSE. See below.
toolsarrayFunction/tool definitions. Claude and GPT both support tool use.
tool_choiceauto|none|namedForce a specific tool or let the model decide.
response_formatobjectJSON output. On Claude a json_schema is enforced and json_object is not — see JSON output for every family.
reasoning_effortstringGPT models only — forwarded as reasoning.effort. Not forwarded for Claude. See Reasoning.
seedintDeterministic sampling where supported.
stopstring|arrayHard stop sequences.
Claude Fable 5.1, Claude Fable 5, Claude Opus 5.5, Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7 and Claude Sonnet 5 do not accept temperature, top_p or top_k. Kunavo removes the three from requests to those models rather than passing the vendor's 400 back to you, so on them the three have no effect.

Streaming

Set stream=True. Kunavo emits server-sent events in OpenAI's exact format: each chunk is a chat.completion.chunk with choices[0].delta.content. The final usage payload arrives with data: [DONE].

stream = client.chat.completions.create(
    model="claude-haiku-4-5",
    messages=[{"role": "user", "content": "Explain B-trees in one paragraph."}],
    stream=True,
)

for chunk in stream:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="", flush=True)
Streaming works for every chat model. For Claude (Anthropic-protocol upstream), we translate the Anthropic event stream into OpenAI deltas on the fly — your SDK doesn't notice the difference.

Tool / function calling

Tool calling works across providers. Define tools as JSON schema; the model returns tool_calls in its message; you execute and feed results back as role: "tool" messages.

tools = [{
    "type": "function",
    "function": {
        "name": "get_current_weather",
        "description": "Get current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {
                "city": {"type": "string"},
                "unit": {"type": "string", "enum": ["c", "f"]},
            },
            "required": ["city"],
        },
    },
}]

resp = client.chat.completions.create(
    model="claude-sonnet-5",
    messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
    tools=tools,
    tool_choice="auto",
)

# Inspect tool calls the model wants to make
for call in resp.choices[0].message.tool_calls or []:
    print(call.function.name, call.function.arguments)

JSON output (response_format)

response_format asks for JSON. What it guarantees depends on the model family, because each family's request is translated to a different upstream mechanism:

Familyjson_schemajson_object
ClaudeEnforced. Sent as Anthropic's structured outputs (output_config.format); the reply is constrained to your schema. Checked on every Claude model Kunavo serves, 2026-09-24.Not enforced. Anthropic has no schema-less JSON mode, so it is accepted and ignored — the reply can open with prose or a code fence.
GPTForwarded as the Responses API's text.format, but the current upstream does not apply it on every request. Validate before you parse.Forwarded as text.format.
import json

resp = client.chat.completions.create(
    model="claude-haiku-4-5",
    messages=[{"role": "user", "content": "Jane Doe <jane@example.com> wants the Pro plan."}],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "signup",
            "schema": {
                "type": "object",
                "properties": {
                    "name": {"type": "string"},
                    "email": {"type": "string"},
                    "plan": {"type": "string", "enum": ["free", "pro", "enterprise"]},
                },
                "required": ["name", "email", "plan"],
                "additionalProperties": False,
            },
        },
    },
)
signup = json.loads(resp.choices[0].message.content)  # constrained to the schema

Four things to know on Claude:

  • The schema goes to Anthropic almost as you wrote it. Anthropic requires additionalProperties: false on objects, so Kunavo adds it to any object that lists its properties and leaves it unset — which is what Pydantic's model_json_schema() produces. What Anthropic still refuses, with a 400 that names the problem: an object left open (additionalProperties set to anything but false, or no properties listed and additionalProperties unset), and oneOf — use anyOf. Anthropic's structured outputs page lists the rest.
  • strict and name are ignored: Anthropic enforces every schema it accepts, strict or not.
  • A schema cannot be combined with a final assistant message (a prefill); Anthropic answers 400.
  • When max_tokens runs out mid-object the JSON is cut off, and finish_reason is "length".

Vision / multimodal input

Models with the vision capability accept image content blocks. Use either an HTTPS URL or a data: base64 URI.

# Pass an image URL or a base64 data URI as part of a multimodal message
resp = client.chat.completions.create(
    model="claude-haiku-4-5",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What's in this image?"},
            {"type": "image_url", "image_url": {
                "url": "https://example.com/cat.jpg"
            }},
        ],
    }],
)
print(resp.choices[0].message.content)

Vision-capable models in the catalog include Claude Opus 5.5 — see the model list for the rest.

For an https image, Kunavo downloads the file itself and hands it to the model inline, because some upstreams ignore a bare image URL. Limits: 16 images per request, 5 MB per image, 20 MB in total, 10 seconds per download. An image that cannot be fetched returns an error naming it rather than being skipped.

Reasoning

What a reasoning control does on this endpoint depends on the family, because each family's request is translated to a different upstream API:

Familyreasoning_effortReasoning tokens in usage
GPTForwarded as the Responses API's reasoning.effort, value for value, on GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, GPT-5.6 Sol, GPT-5.6 Terra and GPT-5.5. The accepted values are the model's own; on GPT-6 Sol and GPT-6 Luna none turns reasoning off and max is the top.Counted inside completion_tokens and broken out as completion_tokens_details.reasoning_tokens, billed at the output rate.
ClaudeNot forwarded. The Anthropic request is built from a fixed list of fields — model, max_tokens, messages, stream, temperature, top_p, top_k, stop, tools, tool_choice and response_format — so reasoning_effort, thinking and anything else outside it are dropped here.No breakdown. Anthropic counts any thinking inside its output_tokens, which is what completion_tokens reports.
# GPT models: reasoning_effort is forwarded as the Responses API's
# reasoning.effort. Reasoning tokens are part of completion_tokens.
resp = client.chat.completions.create(
    model="gpt-6-sol",
    messages=[{"role": "user", "content": "Plan a 3-week MLOps migration."}],
    reasoning_effort="high",
)
details = resp.usage.completion_tokens_details   # absent when no reasoning ran
print(details.reasoning_tokens if details else 0)
Extended thinking on Claude goes through the native Messages API, which passes thinking to Anthropic untranslated.

Prompt caching

A long system prompt, a reference document, a few-shot block — any stable prefix can be cached upstream and replayed on later calls at a fraction of the input price. Cache hits surface in the usage object as prompt_tokens_details.cached_tokens.

GPT models cache automatically — no request change needed. cached_tokens is a subset of prompt_tokens and is billed at a reduced cache-read rate.

Claude caches only the prefix you mark with a cache_control breakpoint. Through this OpenAI-compatible endpoint, attach it to a content block:

# Claude caches prefixes you mark with cache_control. Attach it to a content
# block; later calls reusing that prefix read it at ~10% of the input price.
resp = client.chat.completions.create(
    model="claude-sonnet-5",
    messages=[{
        "role": "user",
        "content": [
            {
                "type": "text",
                "text": LONG_DOCUMENT,            # the stable, reused prefix
                "cache_control": {"type": "ephemeral"},
            },
            {"type": "text", "text": "Summarize the document above."},
        ],
    }],
)
print(resp.usage.prompt_tokens_details.cached_tokens)
For full control over Claude prompt caching — caching the system prompt and tool definitions, multiple breakpoints, and the native cache_creation_input_tokens / cache_read_input_tokens usage fields — use the native Messages API.

Usage object

Every response (and the final streaming chunk) includes a usage object:

FieldMeaning
prompt_tokensInput tokens we billed — cached tokens included.
prompt_tokens_details.cached_tokensCached input — a subset of prompt_tokens, billed at a reduced cache-read rate.
completion_tokensOutput tokens we billed. On GPT models this includes the reasoning tokens; on Claude it is Anthropic's output_tokens.
completion_tokens_details.reasoning_tokensGPT models: the reasoning share of completion_tokens. Absent when zero, and not reported for Claude.
total_tokensprompt_tokens + completion_tokens.

Ollama OpenAI compatibility — /v1/chat/completions

Ollama serves its local models through an OpenAI-compatible API at http://localhost:11434/v1, implementing the same /chat/completions contract documented on this page. That makes the two interchangeable at the client: anything written against Ollama's v1 API runs against Kunavo's by changing the base URL and the model.

from openai import OpenAI

# Before — Ollama's OpenAI-compatible API, local models
client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama",                       # ignored locally
)
resp = client.chat.completions.create(
    model="llama3.2",
    messages=[{"role": "user", "content": "Summarize this changelog."}],
)

# After — the same call against hosted frontier models.
# Two lines changed. Everything below them is untouched.
client = OpenAI(
    base_url="https://api.kunavo.com/v1",   # <- was localhost:11434/v1
    api_key=os.environ["KUNAVO_API_KEY"],
)
resp = client.chat.completions.create(
    model="claude-sonnet-5",  # <- was llama3.2
    messages=[{"role": "user", "content": "Summarize this changelog."}],
)

The pieces that usually break on a provider switch do not break here, because the wire format is the same one:

FeatureOllama v1Kunavo v1
Endpoint path/v1/chat/completions/v1/chat/completions
StreamingSSE, chat.completion.chunkIdentical — same delta shape
Tool callingtools / tool_callsIdentical across Claude and GPT
Vision inputimage_url blocks on multimodal local modelsSame blocks; URL or base64 data URI
api_keyPlaceholder, ignoredA real Kunavo key
modelA pulled tag (llama3.2)A hosted slug (claude-sonnet-5)
usage extrasToken counts onlyAdds cached_tokens and reasoning_tokens
Beyond chatcompletions, embeddingsPlus images, video and audio endpoints
Keeping both? Read base_url, model and the key from environment variables and switch per environment — local models for offline iteration, hosted models where quality matters, with one code path. The full walkthrough, including what Ollama's OpenAI surface does and does not implement, is in the Ollama OpenAI-compatible API guide.

Where are the official docs for Ollama's OpenAI-compatible /v1/chat/completions endpoint?

Ollama documents its OpenAI compatibility layer in its own repository, under docs/openai.md, and that is the authoritative reference for which OpenAI fields the local server implements. This page is the equivalent reference for Kunavo's hosted /v1/chat/completions endpoint: the same OpenAI chat-completions contract, documented field by field above, so you can diff the two surfaces directly. The short version is that the request and response shapes are the same one — Ollama serves it at http://localhost:11434/v1 against locally pulled models, Kunavo serves it at https://api.kunavo.com/v1 against hosted Claude and GPT models.

Is Kunavo's /v1/chat/completions compatible with Ollama's OpenAI API?

Yes — both implement the same OpenAI chat completions contract, so a client written against Ollama's http://localhost:11434/v1 works against https://api.kunavo.com/v1 unchanged. Request fields (messages, temperature, max_tokens, stream, tools, tool_choice, response_format, stop, seed) and response fields (choices[].message, choices[].delta on streams, usage) match. The differences are the ones you would expect from hosted models: the api_key is real rather than a placeholder, model takes a hosted slug such as claude-sonnet-5 instead of a pulled tag like llama3.2, and usage carries cached_tokens and reasoning_tokens that local models do not report.

How do I migrate an Ollama app to hosted Claude or GPT?

Change two lines: set base_url from http://localhost:11434/v1 to https://api.kunavo.com/v1, and set model from the local tag to a hosted slug. Pass a real Kunavo key instead of the placeholder Ollama ignores. Streaming, tool calling and vision code paths need no changes because the SSE and tool_calls shapes are identical.

Can I keep Ollama for development and hosted models in production?

Yes, and it is a common setup — read the base URL, model and key from environment variables and switch per environment. Since both ends speak the same OpenAI protocol there is no second code path to maintain: local Llama or Qwen for offline iteration, hosted Claude or GPT where quality matters. Tool definitions and streaming handlers are shared verbatim.

Which OpenAI endpoints does Kunavo add over Ollama's compatible surface?

Ollama's OpenAI-compatible surface covers /chat/completions, /completions, /embeddings and /models. Kunavo adds the media endpoints on the same key: /v1/images/generations, /v1/images/edits, /v1/video/generations, /v1/videos and /v1/audio/music, plus Claude on the native Anthropic /v1/messages endpoint. Two honest caveats in the other direction: /v1/audio/speech, /v1/audio/transcriptions and /v1/embeddings are implemented as wire formats but currently have no model enabled behind them, so they will not succeed today — and unlike Ollama, Kunavo does not run models locally. GET /v1/models is always the authority on what is callable.

Where to go next