Docs

Prompt caching

Mark a stable prompt prefix once with cache_control — or reach a Claude model in the OpenAI format, which has no such field, and Kunavo places the breakpoints for you. Every subsequent call that reuses the prefix bills it back at a fraction of input price — 10% of input on most chat models, as low as 2.5%. Kunavo's routing layer keeps the cache warm by pinning your conversation to the same upstream key.

Deze documentatie is in het Engels. Een Nederlandstalige gids hebben we nog niet — over betalen lees je hier:Betalen — prijzen in USD, saldo opwaarderen met iDEAL, Bancontact of kaart →

On a 30K-token system prompt with Claude Sonnet 5, a cache read costs $0.14 / 1M instead of $1.4 / 1M. For an agent that re-sends the same context 50 times a session, the cache saves ~90% of the input bill.

How to enable it

Add a cache_control: { type: "ephemeral" } breakpoint on any system block, message, or tool definition. The model caches the span up to that breakpoint on the first call; subsequent calls that share the same prefix bytes hit the cache.

cache_example.py
from openai import OpenAI

client = OpenAI(
    api_key="sk-kn-...",
    base_url="https://api.kunavo.com/v1",
)

# Mark the large, stable prefix once with cache_control.
# Every subsequent call that sends the same prefix reads it back
# from cache at 0.10× the input rate on Claude Sonnet 5.
# Per-model rates are in the pricing table below.
SYSTEM = [
    {
        "type": "text",
        "text": open("long_system_prompt.txt").read(),
        "cache_control": {"type": "ephemeral"},
    },
]

resp = client.chat.completions.create(
    model="claude-sonnet-5",
    messages=[
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": "Summarize the doc above in 3 bullets."},
    ],
)

u = resp.usage
# OpenAI-compat fields populated by Kunavo:
# - prompt_tokens_details.cached_tokens — read from cache
# - cache_creation_input_tokens          — written to cache this call
print(u.prompt_tokens_details.cached_tokens,
      getattr(u, "cache_creation_input_tokens", 0))

Identical wire-format on the native Messages API — cache_control passes through to Claude untranslated:

messages_api.sh
# Native Anthropic Messages API — cache_control passes through untranslated.
curl https://api.kunavo.com/v1/messages \
  -H "Authorization: Bearer $KUNAVO_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-5",
    "system": [
      {
        "type": "text",
        "text": "<your large stable system prompt>",
        "cache_control": { "type": "ephemeral" }
      }
    ],
    "messages": [
      { "role": "user", "content": "Summarize the doc above in 3 bullets." }
    ],
    "max_tokens": 1024
  }'
Stable bytes matter. The cache key is content-addressed. A single different character anywhere in the prefix invalidates the cache for that span. Put dynamic content (timestamps, user names, per-call IDs) after the cached block, not inside it.

Automatic breakpoints on the OpenAI-format endpoints

The OpenAI Chat Completions and Responses formats have no cache_control field at all — OpenAI caches implicitly, on its own servers. Claude does not: it caches only what a breakpoint marks. A client speaking OpenAI therefore has no way to ask for caching, and a straight translation of its request reaches Claude carrying zero breakpoints — so every turn of a long chat re-bills the whole history at the full input rate.

So when you reach a Claude model through /v1/chat/completions or /v1/responses, Kunavo sets the breakpoints the protocol cannot carry. There is nothing to switch on. It applies to Claude models only — every other family is served over a protocol whose vendor caches implicitly, with no breakpoints to place.

no_cache_control.py
# Plain OpenAI Chat Completions — no cache_control anywhere in the body.
# The format has no field for one; Kunavo inserts the breakpoints instead.
resp = client.chat.completions.create(
    model="claude-sonnet-5",
    messages=[
        {"role": "system", "content": LONG_SYSTEM_PROMPT},
        {"role": "user", "content": "first question"},
        {"role": "assistant", "content": "...previous answer..."},
        {"role": "user", "content": "follow-up"},
    ],
)

u = resp.usage
# Turn that writes the cache:
#   u.cache_creation_input_tokens            -> 18,000
#   u.prompt_tokens_details.cached_tokens    ->      0
# Next turn, same prefix, inside the 5-minute window:
#   u.cache_creation_input_tokens            ->    240   (only the new turn)
#   u.prompt_tokens_details.cached_tokens    -> 18,000
#
# Both stuck at 0 on a conversation well over the threshold? The prefix is
# changing between calls — a timestamp or request id inside the cached span.

Where they land, and in what order the budget is spent:

  • A rolling breakpoint on the final message — on the last content block of the last message, so it caches everything before it. Added only once messages already holds an assistant turn: a cache write bills 1.25× input, so marking a prompt nobody re-sends costs a quarter more and saves nothing. A prior assistant turn is the evidence that this exact prefix is about to come back.
  • One on system, one on tools — the stable head of the prompt. Worth marking even on a one-shot call, because a caller that sends a system prompt once almost always sends the same one on the next call.

Two limits bound it. Below 8,192 characters of prompt — messages, system and tool schemas together — nothing is inserted at all: Anthropic silently ignores a breakpoint on a prefix under its own minimum cacheable length (1,024 tokens, 2,048 on Haiku) and charges neither a write nor a read, so one there would buy nothing. And Anthropic accepts at most 4 breakpoints per request — a fifth is a hard 400, not a silent no-op.

Your own breakpoints are never moved or removed. They count against the same budget of 4, and Kunavo only fills positions you left empty. Send your own layout of 4 and nothing is added; send 1 and up to 3 may be added around it. When the budget is tight the deepest position is filled first — a breakpoint late in the prompt caches everything before it, so the rolling one is worth the most. On the native /v1/messages endpoint nothing is inserted at all: your body passes through as sent, and every breakpoint in it is yours.
A breakpoint is not a cache hit. It marks a prefix as cacheable; the hit still needs the same prefix bytes to come back within Anthropic's ephemeral window — five minutes, refreshed on each read — and to land on the same upstream key, which is what the affinity routing below is for. The first call on a fresh prefix is always a write, never a read, and a prompt that changes every turn will not hit no matter where the breakpoints sit.

Nothing about the reporting changes: the same cached_tokens and cache_creation_input_tokens fields described under Response fields are populated whether the breakpoints came from you or from the gateway, so the usage object is how you check that an inserted breakpoint actually earned a hit.

Pricing

Cache rates derive from each model's input price, at the vendor's own ratios — the tables below are the authority, and every rate in them is computed from that model's input price rather than typed in. Cache reads bill at 0.10× input on Anthropic and OpenAI alike, with 2 exceptions: 0.025× on Claude Fable 5.1, 0.05× on Claude Opus 5.5. Cache writes bill at 1.25× input on every Claude model and on GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, GPT-5.6 Sol and GPT-5.6 Terra. Kunavo bills Claude's 1-hour-TTL writes at the same 1.25×, below Anthropic's 2×. Every other model has no separate write price: its writes bill at the plain input rate. Either way, a prefix that is reused even once costs less cached than sent fresh. Every class carries the model's own discount, so its cache prices sit at the same discount off the vendor's list as its input and output.

Anthropic

ModelInputCache readCache write
Claude Fable 5.1$7$0.1750.025×$8.751.25×
Claude Fable 5$7$0.70.10×$8.751.25×
Claude Opus 5.5$2.8$0.140.050×$3.51.25×
Claude Opus 5$3.5$0.350.10×$4.381.25×
Claude Opus 4.8$3.5$0.350.10×$4.381.25×
Claude Opus 4.7$3.5$0.350.10×$4.381.25×
Claude Opus 4.6$3.5$0.350.10×$4.381.25×
Claude Sonnet 5$1.4$0.140.10×$1.751.25×
Claude Sonnet 4.6$2.1$0.210.10×$2.631.25×
Claude Haiku 4.5$0.7$0.070.10×$0.8751.25×

OpenAI

ModelInputCache readCache write
GPT-6 Astra$4$0.40.10×$51.25×
GPT-6 Sol$0.8$0.080.10×$11.25×
GPT-6 Luna$0.04$0.0040.10×$0.051.25×
GPT-5.6 Sol$2$0.20.10×$2.51.25×
GPT-5.6 Terra$0.7$0.070.10×$0.8751.25×
GPT-5.5$2$0.20.10×$21×
Long prompts bill at the vendor's long-context tier. Past the threshold the whole request bills at a multiple of the model's input and output prices, cache reads and writes included: over 272K input tokens on GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, GPT-5.6 Sol, GPT-5.6 Terra and GPT-5.5 (2× input, 1.5× output). Claude has no long-context surcharge: its 1M context bills at the standard rate.
The catalog floor is cache-aware. Kunavo bills max(catalog_cost, upstream × markup) per call — the floor is computed against the cached-input rate, not the fresh rate, so the savings reach you instead of being flattened.
The rates above are catalog prices. The real per-call bill is computed from the upstream credits the model actually consumed (with a markup), then floored at the catalog rate. On cache hits the upstream charges fewer credits — so the upstream-times-markup branch usually drops below catalog and the catalog branch wins. In short: the catalog rate is your worst-case price; cached calls may land below it but never above.

Affinity routing keeps the cache warm

Upstream prompt caches are per-API-key. If every request lands on a random upstream key, the cache hit rate is ~1/N. Kunavo derives a stable hash from your system prompt + first user message and routes all calls with the same prefix to the same upstream key — using weighted rendezvous hashing, so only ~1/N keys are remapped when the pool changes.

You don't need to configure anything. Affinity routing applies automatically to all chat calls. (Image / video / TTS use random routing — there is no per-key cache for those modalities.)

Dashboard metrics

Once your calls start hitting the cache, /app/usage surfaces two rolled-up metrics for the selected window:

  • Hit rate — cached input tokens ÷ total input tokens. Exact, computed from per-call token counts.
  • Saved — USD saved this window, computed per model with its real cache-read rate (no flat-average approximations). This is an estimate at catalog rates — since real bills are floored at catalog and cached calls can land below the floor, the number is a tight upper bound on actual savings. For typical chat workloads the two are within a few percent.

Each call's detail page (/app/usage/<id>) shows the cached + cache-write token counts when they're non-zero, and the actual billed amount.

Response fields

Kunavo surfaces cache tokens in both response shapes. On the OpenAI-compatible /v1/chat/completions endpoint:

FieldMeaning
usage.prompt_tokensTotal input tokens (includes cached + cache-write).
usage.prompt_tokens_details.cached_tokensSubset of input served from cache.
usage.cache_creation_input_tokensTokens written to cache this call (billed at the model's cache-write rate, in the tables above).

On the native Messages API at /v1/messages, the Anthropic-original fields pass through:

FieldMeaning
input_tokensFresh (uncached) input.
cache_read_input_tokensServed from cache, billed 0.10× input (0.025× on Claude Fable 5.1, 0.05× on Claude Opus 5.5).
cache_creation_input_tokensWritten to cache this call (billed 1.25× input — 1-hour-TTL writes too, below Anthropic's 2×).

Where to go next

  • Messages API reference — native Anthropic shape with worked cache examples.
  • Claude prompt caching — the usage-object numbers that diagnose a cache, and the three ways routing through a gateway breaks a hit rate.
  • Billing & the ledger — how the cache-aware max() floor interacts with upstream cost.
  • Full pricing table — fresh input / output prices for every chat model, with the cache sub-table at the bottom.