Docs

Messages API

The native Anthropic Messages API for the Claude family. Send the real Anthropic request shape — cache_control, tools, and extended thinking pass straight through, untranslated.

Deze documentatie is in het Engels. Een Nederlandstalige gids hebben we nog niet — over betalen lees je hier:Betalen — prijzen in USD, saldo opwaarderen met iDEAL, Bancontact of kaart →

Endpoint: POST /v1/messages. Request and response match Anthropic's Messages API exactly, so the Anthropic SDK — or any tool that speaks it — works against Kunavo by changing only the base URL and the key.

Which endpoint should I use?

Kunavo exposes the Claude family two ways. They bill identically; pick by which request shape your code already speaks.

EndpointShapeUse it when
/v1/chat/completionsOpenAI chatYou already use the OpenAI SDK, or you want one code path across Claude and GPT.
/v1/messagesAnthropic MessagesYou use the Anthropic SDK / Claude Code, or you want native cache_control, tools and thinking with no translation layer.

Authentication

Point the base URL at https://api.kunavo.com — the Anthropic SDK appends /v1/messages itself. Pass your Kunavo key (sk-kn-...) as the API key; the endpoint accepts it via either the x-api-key header (Anthropic SDK default) or Authorization: Bearer.

from anthropic import Anthropic

client = Anthropic(
    api_key="sk-kn-...",                  # your Kunavo key
    base_url="https://api.kunavo.com",    # SDK appends /v1/messages
)

msg = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    system="You are a senior staff engineer.",
    messages=[
        {"role": "user", "content": "Pros and cons of event sourcing?"},
    ],
)
print(msg.content[0].text)

Prompt caching

This is the main reason to use the native endpoint. Attach a cache_control breakpoint to any system block, message content block, or tool definition. The marked prefix is written to cache once, then replayed on later calls that share it.

# Mark a large, stable prefix with cache_control. The first call writes that
# span to cache at 1.25× the input token price; later calls that reuse it read
# it back at ~10%.
msg = client.messages.create(
    model="claude-opus-4-7",
    max_tokens=1024,
    system=[
        {
            "type": "text",
            "text": LONG_STYLE_GUIDE,            # tens of thousands of tokens
            "cache_control": {"type": "ephemeral"},
        },
    ],
    messages=[{"role": "user", "content": "Review this PR against the guide."}],
)
u = msg.usage
print(u.cache_creation_input_tokens, u.cache_read_input_tokens)

Pricing for the Claude family on Kunavo:

  • cache_read_input_tokens — served from cache, billed at 0.10× the input rate (a 90% discount); 0.025× on Claude Fable 5.1, 0.05× on Claude Opus 5.5.
  • cache_creation_input_tokens — written to cache, billed at 1.25× the input rate. Kunavo bills 1-hour-TTL writes at the same 1.25×, below Anthropic's 2×.
Kunavo keeps your cache warm automatically: requests that share a prompt prefix are routed to the same upstream account, with no configuration on your side.

Streaming

Set stream: true (or use the SDK stream() helper). Kunavo forwards Anthropic's native event stream — message_start, content_block_delta, message_delta, message_stop — through verbatim.

with client.messages.stream(
    model="claude-sonnet-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Explain CRDTs in one paragraph."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

Images

Send images as base64 sources or url sources. A url source is the one part of the body Kunavo rewrites: it downloads an https image itself (or unpacks a data: URI) and forwards it as base64, because the upstream ignores a url source and would answer as though no image had been sent. Images inside a tool_result are handled the same way. Limits: 16 images per request, 5 MB per image, 20 MB in total; an image that cannot be fetched returns an error rather than being skipped.

Usage object

Every response (and the streaming message_delta / message_start events) carries a native Anthropic usage object:

FieldMeaning
input_tokensFresh (uncached) input tokens.
cache_read_input_tokensTokens served from cache, billed 0.10× input (0.025× on Claude Fable 5.1, 0.05× on Claude Opus 5.5).
cache_creation_input_tokensTokens written to cache this call (billed 1.25× input).
output_tokensOutput tokens, including extended thinking.

Server tools

Anthropic's server tools (web_search, web_fetch, code_execution) go through as ordinary tool definitions. What they add to the conversation (search results, a fetched page, program output) comes back in usage as tokens and bills at the rates above.

Web search also bills per search, on top of those tokens: Anthropic lists it at $10 per 1,000 searches, and Kunavo charges that at each model's own discount. The count is usage.server_tool_use.web_search_requests, in the response or the final message_delta of a stream; a search that returns an error (such as max_uses_exceeded) is not counted. Web fetch and code execution cost nothing beyond their tokens.

Per 1,000 web searchesModels
$7.00claude-fable-5-1, claude-fable-5, claude-opus-5-5, claude-opus-5, claude-opus-4-8, claude-opus-4-7, claude-opus-4-6, claude-sonnet-5, claude-sonnet-4-6, claude-haiku-4-5

Supported models

Any enabled Claude model accepts /v1/messages. Send the Kunavo slug as model:

  • claude-fable-5, claude-sonnet-5
  • claude-opus-5, claude-opus-4-7, claude-opus-4-6
  • claude-sonnet-4-6, claude-haiku-4-5

GPT models are not available here — call them via /v1/chat/completions.

Raw HTTP

No SDK required — the endpoint is a plain JSON POST:

curl https://api.kunavo.com/v1/messages \
  -H "x-api-key: sk-kn-..." \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-5",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Hello, Claude"}]
  }'

Where to go next