Docs
Prompt caching
Mark a stable prompt prefix once with cache_control — or reach a Claude model in the OpenAI format, which has no such field, and Kunavo places the breakpoints for you. Every subsequent call that reuses the prefix bills it back at a fraction of input price — 10% of input on most chat models, as low as 2.5%. Kunavo's routing layer keeps the cache warm by pinning your conversation to the same upstream key.
On a 30K-token system prompt with Claude Sonnet 5, a cache read costs $0.14 / 1M instead of $1.4 / 1M. For an agent that re-sends the same context 50 times a session, the cache saves ~90% of the input bill.
How to enable it
Add a cache_control: { type: "ephemeral" } breakpoint on any system block, message, or tool definition. The model caches the span up to that breakpoint on the first call; subsequent calls that share the same prefix bytes hit the cache.
from openai import OpenAI
client = OpenAI(
api_key="sk-kn-...",
base_url="https://api.kunavo.com/v1",
)
# Mark the large, stable prefix once with cache_control.
# Every subsequent call that sends the same prefix reads it back
# from cache at 0.10× the input rate on Claude Sonnet 5.
# Per-model rates are in the pricing table below.
SYSTEM = [
{
"type": "text",
"text": open("long_system_prompt.txt").read(),
"cache_control": {"type": "ephemeral"},
},
]
resp = client.chat.completions.create(
model="claude-sonnet-5",
messages=[
{"role": "system", "content": SYSTEM},
{"role": "user", "content": "Summarize the doc above in 3 bullets."},
],
)
u = resp.usage
# OpenAI-compat fields populated by Kunavo:
# - prompt_tokens_details.cached_tokens — read from cache
# - cache_creation_input_tokens — written to cache this call
print(u.prompt_tokens_details.cached_tokens,
getattr(u, "cache_creation_input_tokens", 0))Identical wire-format on the native Messages API — cache_control passes through to Claude untranslated:
# Native Anthropic Messages API — cache_control passes through untranslated.
curl https://api.kunavo.com/v1/messages \
-H "Authorization: Bearer $KUNAVO_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5",
"system": [
{
"type": "text",
"text": "<your large stable system prompt>",
"cache_control": { "type": "ephemeral" }
}
],
"messages": [
{ "role": "user", "content": "Summarize the doc above in 3 bullets." }
],
"max_tokens": 1024
}'Automatic breakpoints on the OpenAI-format endpoints
The OpenAI Chat Completions and Responses formats have no cache_control field at all — OpenAI caches implicitly, on its own servers. Claude does not: it caches only what a breakpoint marks. A client speaking OpenAI therefore has no way to ask for caching, and a straight translation of its request reaches Claude carrying zero breakpoints — so every turn of a long chat re-bills the whole history at the full input rate.
So when you reach a Claude model through /v1/chat/completions or /v1/responses, Kunavo sets the breakpoints the protocol cannot carry. There is nothing to switch on. It applies to Claude models only — every other family is served over a protocol whose vendor caches implicitly, with no breakpoints to place.
# Plain OpenAI Chat Completions — no cache_control anywhere in the body.
# The format has no field for one; Kunavo inserts the breakpoints instead.
resp = client.chat.completions.create(
model="claude-sonnet-5",
messages=[
{"role": "system", "content": LONG_SYSTEM_PROMPT},
{"role": "user", "content": "first question"},
{"role": "assistant", "content": "...previous answer..."},
{"role": "user", "content": "follow-up"},
],
)
u = resp.usage
# Turn that writes the cache:
# u.cache_creation_input_tokens -> 18,000
# u.prompt_tokens_details.cached_tokens -> 0
# Next turn, same prefix, inside the 5-minute window:
# u.cache_creation_input_tokens -> 240 (only the new turn)
# u.prompt_tokens_details.cached_tokens -> 18,000
#
# Both stuck at 0 on a conversation well over the threshold? The prefix is
# changing between calls — a timestamp or request id inside the cached span.Where they land, and in what order the budget is spent:
- A rolling breakpoint on the final message — on the last content block of the last message, so it caches everything before it. Added only once
messagesalready holds an assistant turn: a cache write bills1.25×input, so marking a prompt nobody re-sends costs a quarter more and saves nothing. A prior assistant turn is the evidence that this exact prefix is about to come back. - One on
system, one ontools— the stable head of the prompt. Worth marking even on a one-shot call, because a caller that sends a system prompt once almost always sends the same one on the next call.
Two limits bound it. Below 8,192 characters of prompt — messages, system and tool schemas together — nothing is inserted at all: Anthropic silently ignores a breakpoint on a prefix under its own minimum cacheable length (1,024 tokens, 2,048 on Haiku) and charges neither a write nor a read, so one there would buy nothing. And Anthropic accepts at most 4 breakpoints per request — a fifth is a hard 400, not a silent no-op.
/v1/messages endpoint nothing is inserted at all: your body passes through as sent, and every breakpoint in it is yours.Nothing about the reporting changes: the same cached_tokens and cache_creation_input_tokens fields described under Response fields are populated whether the breakpoints came from you or from the gateway, so the usage object is how you check that an inserted breakpoint actually earned a hit.
Pricing
Cache rates derive from each model's input price, at the vendor's own ratios — the tables below are the authority, and every rate in them is computed from that model's input price rather than typed in. Cache reads bill at 0.10× input on Anthropic and OpenAI alike, with 2 exceptions: 0.025× on Claude Fable 5.1, 0.05× on Claude Opus 5.5. Cache writes bill at 1.25× input on every Claude model and on GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, GPT-5.6 Sol and GPT-5.6 Terra. Kunavo bills Claude's 1-hour-TTL writes at the same 1.25×, below Anthropic's 2×. Every other model has no separate write price: its writes bill at the plain input rate. Either way, a prefix that is reused even once costs less cached than sent fresh. Every class carries the model's own discount, so its cache prices sit at the same discount off the vendor's list as its input and output.
Anthropic
| Model | Input | Cache read | Cache write |
|---|---|---|---|
| Claude Fable 5.1 | $7 | $0.1750.025× | $8.751.25× |
| Claude Fable 5 | $7 | $0.70.10× | $8.751.25× |
| Claude Opus 5.5 | $2.8 | $0.140.050× | $3.51.25× |
| Claude Opus 5 | $3.5 | $0.350.10× | $4.381.25× |
| Claude Opus 4.8 | $3.5 | $0.350.10× | $4.381.25× |
| Claude Opus 4.7 | $3.5 | $0.350.10× | $4.381.25× |
| Claude Opus 4.6 | $3.5 | $0.350.10× | $4.381.25× |
| Claude Sonnet 5 | $1.4 | $0.140.10× | $1.751.25× |
| Claude Sonnet 4.6 | $2.1 | $0.210.10× | $2.631.25× |
| Claude Haiku 4.5 | $0.7 | $0.070.10× | $0.8751.25× |
OpenAI
| Model | Input | Cache read | Cache write |
|---|---|---|---|
| GPT-6 Astra | $4 | $0.40.10× | $51.25× |
| GPT-6 Sol | $0.8 | $0.080.10× | $11.25× |
| GPT-6 Luna | $0.04 | $0.0040.10× | $0.051.25× |
| GPT-5.6 Sol | $2 | $0.20.10× | $2.51.25× |
| GPT-5.6 Terra | $0.7 | $0.070.10× | $0.8751.25× |
| GPT-5.5 | $2 | $0.20.10× | $21× |
max(catalog_cost, upstream × markup) per call — the floor is computed against the cached-input rate, not the fresh rate, so the savings reach you instead of being flattened.Affinity routing keeps the cache warm
Upstream prompt caches are per-API-key. If every request lands on a random upstream key, the cache hit rate is ~1/N. Kunavo derives a stable hash from your system prompt + first user message and routes all calls with the same prefix to the same upstream key — using weighted rendezvous hashing, so only ~1/N keys are remapped when the pool changes.
You don't need to configure anything. Affinity routing applies automatically to all chat calls. (Image / video / TTS use random routing — there is no per-key cache for those modalities.)
Dashboard metrics
Once your calls start hitting the cache, /app/usage surfaces two rolled-up metrics for the selected window:
- Hit rate — cached input tokens ÷ total input tokens. Exact, computed from per-call token counts.
- Saved — USD saved this window, computed per model with its real cache-read rate (no flat-average approximations). This is an estimate at catalog rates — since real bills are floored at catalog and cached calls can land below the floor, the number is a tight upper bound on actual savings. For typical chat workloads the two are within a few percent.
Each call's detail page (/app/usage/<id>) shows the cached + cache-write token counts when they're non-zero, and the actual billed amount.
Response fields
Kunavo surfaces cache tokens in both response shapes. On the OpenAI-compatible /v1/chat/completions endpoint:
| Field | Meaning |
|---|---|
usage.prompt_tokens | Total input tokens (includes cached + cache-write). |
usage.prompt_tokens_details.cached_tokens | Subset of input served from cache. |
usage.cache_creation_input_tokens | Tokens written to cache this call (billed at the model's cache-write rate, in the tables above). |
On the native Messages API at /v1/messages, the Anthropic-original fields pass through:
| Field | Meaning |
|---|---|
input_tokens | Fresh (uncached) input. |
cache_read_input_tokens | Served from cache, billed 0.10× input (0.025× on Claude Fable 5.1, 0.05× on Claude Opus 5.5). |
cache_creation_input_tokens | Written to cache this call (billed 1.25× input — 1-hour-TTL writes too, below Anthropic's 2×). |
Where to go next
- Messages API reference — native Anthropic shape with worked cache examples.
- Claude prompt caching — the usage-object numbers that diagnose a cache, and the three ways routing through a gateway breaks a hit rate.
- Billing & the ledger — how the cache-aware max() floor interacts with upstream cost.
- Full pricing table — fresh input / output prices for every chat model, with the cache sub-table at the bottom.