Docs
Messages API
The native Anthropic Messages API for the Claude family. Send the real Anthropic request shape — cache_control, tools, and extended thinking pass straight through, untranslated.
Endpoint: POST /v1/messages. Request and response match Anthropic's Messages API exactly, so the Anthropic SDK — or any tool that speaks it — works against Kunavo by changing only the base URL and the key.
Which endpoint should I use?
Kunavo exposes the Claude family two ways. They bill identically; pick by which request shape your code already speaks.
| Endpoint | Shape | Use it when |
|---|---|---|
/v1/chat/completions | OpenAI chat | You already use the OpenAI SDK, or you want one code path across Claude and GPT. |
/v1/messages | Anthropic Messages | You use the Anthropic SDK / Claude Code, or you want native cache_control, tools and thinking with no translation layer. |
Authentication
Point the base URL at https://api.kunavo.com — the Anthropic SDK appends /v1/messages itself. Pass your Kunavo key (sk-kn-...) as the API key; the endpoint accepts it via either the x-api-key header (Anthropic SDK default) or Authorization: Bearer.
from anthropic import Anthropic
client = Anthropic(
api_key="sk-kn-...", # your Kunavo key
base_url="https://api.kunavo.com", # SDK appends /v1/messages
)
msg = client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
system="You are a senior staff engineer.",
messages=[
{"role": "user", "content": "Pros and cons of event sourcing?"},
],
)
print(msg.content[0].text)Prompt caching
This is the main reason to use the native endpoint. Attach a cache_control breakpoint to any system block, message content block, or tool definition. The marked prefix is written to cache once, then replayed on later calls that share it.
# Mark a large, stable prefix with cache_control. The first call writes that
# span to cache at 1.25× the input token price; later calls that reuse it read
# it back at ~10%.
msg = client.messages.create(
model="claude-opus-4-7",
max_tokens=1024,
system=[
{
"type": "text",
"text": LONG_STYLE_GUIDE, # tens of thousands of tokens
"cache_control": {"type": "ephemeral"},
},
],
messages=[{"role": "user", "content": "Review this PR against the guide."}],
)
u = msg.usage
print(u.cache_creation_input_tokens, u.cache_read_input_tokens)Pricing for the Claude family on Kunavo:
cache_read_input_tokens— served from cache, billed at 0.10× the input rate (a 90% discount); 0.025× on Claude Fable 5.1, 0.05× on Claude Opus 5.5.cache_creation_input_tokens— written to cache, billed at 1.25× the input rate. Kunavo bills 1-hour-TTL writes at the same 1.25×, below Anthropic's 2×.
Streaming
Set stream: true (or use the SDK stream() helper). Kunavo forwards Anthropic's native event stream — message_start, content_block_delta, message_delta, message_stop — through verbatim.
with client.messages.stream(
model="claude-sonnet-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Explain CRDTs in one paragraph."}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)Images
Send images as base64 sources or url sources. A url source is the one part of the body Kunavo rewrites: it downloads an https image itself (or unpacks a data: URI) and forwards it as base64, because the upstream ignores a url source and would answer as though no image had been sent. Images inside a tool_result are handled the same way. Limits: 16 images per request, 5 MB per image, 20 MB in total; an image that cannot be fetched returns an error rather than being skipped.
Usage object
Every response (and the streaming message_delta / message_start events) carries a native Anthropic usage object:
| Field | Meaning |
|---|---|
input_tokens | Fresh (uncached) input tokens. |
cache_read_input_tokens | Tokens served from cache, billed 0.10× input (0.025× on Claude Fable 5.1, 0.05× on Claude Opus 5.5). |
cache_creation_input_tokens | Tokens written to cache this call (billed 1.25× input). |
output_tokens | Output tokens, including extended thinking. |
Server tools
Anthropic's server tools (web_search, web_fetch, code_execution) go through as ordinary tool definitions. What they add to the conversation (search results, a fetched page, program output) comes back in usage as tokens and bills at the rates above.
Web search also bills per search, on top of those tokens: Anthropic lists it at $10 per 1,000 searches, and Kunavo charges that at each model's own discount. The count is usage.server_tool_use.web_search_requests, in the response or the final message_delta of a stream; a search that returns an error (such as max_uses_exceeded) is not counted. Web fetch and code execution cost nothing beyond their tokens.
| Per 1,000 web searches | Models |
|---|---|
| $7.00 | claude-fable-5-1, claude-fable-5, claude-opus-5-5, claude-opus-5, claude-opus-4-8, claude-opus-4-7, claude-opus-4-6, claude-sonnet-5, claude-sonnet-4-6, claude-haiku-4-5 |
Supported models
Any enabled Claude model accepts /v1/messages. Send the Kunavo slug as model:
claude-fable-5,claude-sonnet-5claude-opus-5,claude-opus-4-7,claude-opus-4-6claude-sonnet-4-6,claude-haiku-4-5
GPT models are not available here — call them via /v1/chat/completions.
Raw HTTP
No SDK required — the endpoint is a plain JSON POST:
curl https://api.kunavo.com/v1/messages \
-H "x-api-key: sk-kn-..." \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Hello, Claude"}]
}'