Docs
OpenClaw
OpenClaw reaches any endpoint through one models.providers entry. For an agent that never stops, the entry is the short part: this page also covers which side places the cache breakpoints on each wire, what a day of heartbeats costs, and what a 402 does to the gateway.
One models.providers entry in ~/.openclaw/openclaw.json — baseUrl https://api.kunavo.com, api "anthropic-messages" — puts an always-on OpenClaw agent on Claude; cacheRetention is set beside it, because a custom Anthropic endpoint gets no cache markers until it is.
// ~/.openclaw/openclaw.json — merge into the file you already have
{
models: {
mode: "merge",
providers: {
kunavo: {
baseUrl: "https://api.kunavo.com", // origin — no /v1 on this wire
apiKey: "${KUNAVO_API_KEY}", // from the environment or ~/.openclaw/.env
api: "anthropic-messages",
models: [
{
id: "claude-sonnet-5",
name: "Claude Sonnet 5",
reasoning: true,
input: ["text", "image"],
contextWindow: 1000000,
contextTokens: 200000, // optional: compact here, not at 1M
maxTokens: 32000,
},
{
id: "claude-haiku-4-5",
name: "Claude Haiku 4.5",
input: ["text", "image"],
contextWindow: 200000,
maxTokens: 16000,
},
],
},
},
},
agents: {
defaults: {
model: { primary: "kunavo/claude-sonnet-5" },
models: {
// Required for caching: a custom Anthropic endpoint gets no cache
// markers from OpenClaw until cacheRetention is set explicitly.
"kunavo/claude-sonnet-5": { params: { cacheRetention: "short" } },
"kunavo/claude-haiku-4-5": { params: { cacheRetention: "short" } },
},
},
},
}https://api.kunavo.com, no /v1. OpenClaw's own example of an Anthropic-compatible provider says the base URL should omit /v1, because the Anthropic client appends it. The OpenAI-compatible wire, further down, is the one that keeps the suffix.cacheRetention lines are what turn prompt caching on. For a custom Anthropic endpoint OpenClaw sends cache markers only when cacheRetention is set explicitly, and Kunavo's /v1/messages adds none of its own. Leave them out and every turn bills the whole conversation again as fresh input.maxTokens is the output limit OpenClaw works within for a model, and on Kunavo the output cap of a request is part of what it reserves against your balance before it runs. The catalog allows up to 128,000 output tokens on Claude Sonnet 5; the smaller figure in the block is ample for an agent turn and keeps that reservation low.contextTokens is optional. Claude Sonnet 5 has a 1,000,000-token window at a flat rate, and a session that never ends will grow into it; contextTokens gives OpenClaw a smaller working budget, so it compacts long before every turn re-sends the whole window.sk-kn-) and add credit from $10 — calls are paid from that balance, and failed calls are not billed. The dashboard then opens on the OpenClaw setup.Step by step
- Create a key at
/app/keysand copy it — it is shown once. - Give the Gateway the key: add
KUNAVO_API_KEY=sk-kn-...to~/.openclaw/.env, or export it in the environment the Gateway starts in. The${KUNAVO_API_KEY}in the block is replaced from there when the config loads. - Merge the block into
~/.openclaw/openclaw.json, keeping the providers, agents and channels you already have. The file is JSON5, so the comments can stay. - Run
openclaw config validate. OpenClaw refuses to start when the file holds a setting it does not recognise, so a typo is better found here than at the next restart. - Run
openclaw models list --provider kunavoand check that both ids are listed. If a running Gateway has not picked the change up,openclaw gateway restart. - Open a fresh session with
/new— an existing one keeps the model it already had — send two messages, then read/usage tokens: the second turn should reportcacheRead.
Checked against OpenClaw's custom providers reference on October 5, 2026. Third-party settings move; if a field name here no longer matches what you see, that page is the authority, not this one.
Verify before you debug the client
One request settles whether a failure is the endpoint, the key, or the configuration file. If this returns JSON, the same base URL and key work in OpenClaw.
# Settles whether a failure is the endpoint, the key, or the client.
curl -sS https://api.kunavo.com/v1/messages \
-H "Authorization: Bearer sk-kn-..." \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model":"claude-sonnet-5","max_tokens":16,"messages":[{"role":"user","content":"ping"}]}'Which model id to put in the field
Every text model is reachable as a model id — the live list is GET /v1/models, and the catalog with prices is on the models page. Rates are USD per 1M tokens, input / output.
| Model id | Kunavo in / out | Where it fits in OpenClaw |
|---|---|---|
claude-sonnet-5 | $1.40 / $7.00 | the main agent — tool loops and everyday requests |
claude-opus-5-5 | $2.80 / $14.00 | the step up for long or difficult tasks; add it as another row and switch with /model |
claude-haiku-4-5 | $0.70 / $3.50 | heartbeats, session titles and other short background turns |
claude-fable-5 | $7.00 / $35.00 | the top tier — price a day on it with the heartbeat table below before leaving an agent on it |
The OpenAI-compatible wire
The same key reaches every other model family through /v1/chat/completions. Register it as a second provider entry, so the two wires stay apart, and refer to its models as kunavo-openai/<id>:
// ~/.openclaw/openclaw.json — a second entry, beside "kunavo"
{
models: {
providers: {
"kunavo-openai": {
baseUrl: "https://api.kunavo.com/v1", // this wire keeps /v1
apiKey: "${KUNAVO_API_KEY}",
api: "openai-completions",
models: [
{
id: "gpt-6-sol",
name: "GPT-6 Sol",
reasoning: true,
input: ["text"],
contextWindow: 1050000,
maxTokens: 32000,
},
],
},
},
},
}
// then: /model kunavo-openai/gpt-6-solThree things differ from the block at the top. The base URL keeps /v1, the form OpenClaw's own custom-provider examples use. api is openai-completions — also what OpenClaw assumes when a custom provider gives a baseUrl and no api. And maxTokens stops being optional in practice: when a model's output limit is unknown, OpenClaw sends no cap at all on this wire, and Kunavo then ends a Claude reply at 4,096 tokens.
Claude ids work here as well, and this is the wire to use if you want one entry for everything. Two things change for them: thinking levels are not forwarded for Claude on chat completions, and the cache breakpoints are placed by Kunavo, not by OpenClaw.
Prompt caching on each wire
On the Anthropic wire OpenClaw places the cache breakpoints itself, but for a custom endpoint only when cacheRetention is set. Its prompt caching reference is specific about it: the default of short is seeded for the anthropic and anthropic-vertex providers only, and every other Anthropic-family route needs an explicit value. Kunavo's /v1/messages forwards the body as it was sent and adds no breakpoint, so a config without those lines caches nothing.
short asks for the five-minute entry and long for a one-hour one. Kunavo forwards either marker and bills the write at the same rate. Whether a one-hour entry is still there when the next heartbeat arrives is worth confirming on your own usage before you plan a cadence around it: a turn that reports cacheRead kept its cache, and one that reports cacheWrite again did not. /usage tokens and /status show both counters.
On the OpenAI-compatible wire it is the other way round. OpenClaw sends no cache hints to a proxy endpoint, and Kunavo places the breakpoints for Claude models itself — on the system prompt, the tool definitions and the end of the conversation — once the prompt is long enough to cache. There is nothing to configure, and GPT models are cached implicitly by their vendor.
Whichever side places the breakpoints, the bill reads the same. On Claude Sonnet 5 a cache read is $0.14 per 1M tokens against $1.40 for fresh input, and a cache write is $1.75 — the write premium Claude carries over input, charged at that same rate when the entry asks for a one-hour lifetime. An entry lasts five minutes and every read renews it, so what an agent pays depends less on the model than on whether its next request arrives inside that window. Every model's cache rates are on the prompt caching page.
OpenClaw can show the same arithmetic locally. Its /usage cost summary and the cost line in /status need a cost object on each model row, and without one they read zero while Kunavo bills as usual. These rows are generated from the live catalog:
// merge into the rows of models.providers.kunavo.models — USD per 1M tokens
{ id: "claude-sonnet-5", cost: { input: 1.4, output: 7, cacheRead: 0.14, cacheWrite: 1.75 } },
{ id: "claude-haiku-4-5", cost: { input: 0.7, output: 3.5, cacheRead: 0.07, cacheWrite: 0.875 } },
{ id: "claude-opus-5-5", cost: { input: 2.8, output: 14, cacheRead: 0.14, cacheWrite: 3.5 } },
{ id: "claude-fable-5", cost: { input: 7, output: 35, cacheRead: 0.7, cacheWrite: 8.75 } },What an always-on agent costs per day
An OpenClaw agent bills while nobody is talking to it, because of the heartbeat: a scheduled agent turn that runs every 30 minutes by default, which is 48 a day. Unless told otherwise it runs in the main session and re-sends the conversation — OpenClaw's reference puts such a run at about 100,000 tokens, and at a few thousand once it is isolated. Thirty minutes is longer than the five-minute cache window, so each run bills its whole prompt again: at the input rate, or at the higher write rate where a breakpoint is placed. The table prices one idle day at the input rate, with 5,000 tokens for the isolated run:
| Model on the heartbeat | Input rate per 1M tokens | 48 runs in the main session | 48 isolated runs |
|---|---|---|---|
claude-haiku-4-5 | $0.70 | $3.36 | $0.17 |
claude-sonnet-5 | $1.40 | $6.72 | $0.34 |
claude-opus-5-5 | $2.80 | $13.44 | $0.67 |
claude-fable-5 | $7.00 | $33.60 | $1.68 |
The block below is the cheap corner of that table: heartbeats on Haiku, in an isolated session, without the workspace bootstrap files, and only during waking hours. Every setting in it is from OpenClaw's heartbeat reference. A longer every is the other lever, and "0m" turns the recurring run off.
// ~/.openclaw/openclaw.json — what decides the cost of an idle day
{
agents: {
defaults: {
utilityModel: "kunavo/claude-haiku-4-5", // titles and other short internal tasks
heartbeat: {
every: "30m", // the default with an API key
model: "kunavo/claude-haiku-4-5", // wake-ups on the cheapest tier
isolatedSession: true, // a fresh session, not the whole conversation
lightContext: true, // skip the workspace bootstrap files
activeHours: { start: "08:00", end: "24:00" },
},
},
},
}model and isolatedSession together. OpenClaw's heartbeat page warns that a heartbeat which switches a shared session to a smaller model can leave that model in place for the next real turn; a fresh session per run avoids it.The hours when the agent is actually working are the other half of the bill, and there the cache decides it. Take 100 consecutive requests, each re-sending a 100,000-token context with 2,000 new tokens on top and returning 800 tokens of output. On Claude Sonnet 5 that is about $2.31 while the context is read from cache, and about $14.84 when every request bills it as fresh input. Same work, same model: the difference is whether the breakpoints are there and the requests are less than five minutes apart.
For scale, measured and not assumed: across the Kunavo accounts that run an always-on agent, a median active day has cost $12.67 and a 90th-percentile day about $163. Those are amounts billed up to October 5, 2026, at the rates in force on each day. It is a small group, so read it as the width of the range, not as a forecast for your agent.
When the balance runs out
Kunavo is prepaid: every call is paid from the wallet, and an agent that works while you sleep empties it while you sleep. A request the wallet cannot cover is refused with HTTP 402 and the code insufficient_balance, on either wire, and nothing is charged for it. The refusal comes before the wallet reads zero: each request first reserves its worst-case cost, its prompt plus the largest reply it is allowed to produce, so the larger the output cap an agent asks for, the earlier its calls start to bounce. The error says how far short it was, in balance_usd and needed_usd.
OpenClaw decides what a 402 means from its message. By the rules in OpenClaw 2026.9.8, the wallet refusal is a billing failure, and its failover reference says what follows: the credential is disabled for ten minutes, the run moves to the next model in agents.defaults.model.fallbacks, and recharging does not clear the window — so after a top-up the agent can stay off kunavo/… until it ends. The refusal for a key's monthly limit is read differently. Its message names a limit that resets, which the same rules treat as a rate limit: OpenClaw retries, then cools the credential down for 30 seconds at first and five minutes at most. openclaw models status lists a disabled credential and when it recovers.
Two settings keep an unattended agent out of that state, and they do different jobs:
- Auto-recharge, under Billing. Save a card once and set three numbers: the balance below which to top up, the amount to add each time, and a monthly cap. The wallet then refills within seconds of a call that takes it under the threshold. A request that arrives while the wallet is still short waits for that charge and is then served instead of refused. A
402still comes back when the charge cannot be made — a declined card, the monthly cap reached — or when one request reserves more than the wallet holds after the top-up. It needs a card or Link — Alipay, WeChat Pay, Pix and the other local methods cannot be charged automatically. - A monthly limit on the key, under API Keys. Give the agent a key of its own and set the most that key may spend in a calendar month. Past that figure its calls are refused with a
402and nothing is charged, while your other keys keep working. That is the ceiling a runaway loop needs, and one the wallet cannot provide, because every key draws on the same wallet.
Set the recharge threshold above what one request reserves, and size the amount from a day of your agent, not from the minimum: the smallest top-up is $10, and the median always-on day above is $12.67. The limits on auto-recharge are on the billing page, and the whole error body is on the errors page.
Related guides
- Best API for OpenClaw — choosing a provider and a model for each kind of task.
- OpenClaw pricing — the whole operating bill: software, hosting, models and tools.
- OpenClaw with multiple agents and models — routing each agent to its own model and attributing the bill by route.
- Hermes vs OpenClaw — and the same setup for the other agent, on the Hermes Agent page.
Frequently asked questions
How do I add a custom provider to OpenClaw?
Add an entry under models.providers in ~/.openclaw/openclaw.json, keyed by a provider id of your choosing. It takes a baseUrl, an apiKey (usually a ${ENV_VAR} reference), an api type — openai-completions, openai-responses or anthropic-messages among others — and a models array whose entries need at least an id. Then point agents.defaults.model.primary at provider-id/model-id. OpenClaw validates the file strictly, so run openclaw config validate before restarting the Gateway.
Does the OpenClaw base URL need /v1?
It depends on the api type. With api "anthropic-messages" the base URL is the bare origin, because the Anthropic client appends /v1/messages itself — for Kunavo, https://api.kunavo.com. With api "openai-completions" it keeps the suffix, the form OpenClaw's own custom-provider examples use — for Kunavo, https://api.kunavo.com/v1. The wrong form for the wire is the usual reason an endpoint that is up answers 404.
Does prompt caching work in OpenClaw through a custom endpoint?
Yes, and which side does the work depends on the wire. On a custom anthropic-messages endpoint OpenClaw sends cache markers only when cacheRetention is set explicitly — short for a five-minute entry, long for a one-hour one — so the setting belongs in agents.defaults.models for each model you use. On an OpenAI-compatible endpoint OpenClaw sends no cache hints to a proxy, and Kunavo places the breakpoints for Claude models itself. Either way, cacheRead and cacheWrite in /usage tokens show whether it is working.
How much does it cost to run OpenClaw all day?
Price the heartbeat first, because it runs whether or not anyone talks to the agent. At OpenClaw's default of one heartbeat every 30 minutes a day holds 48 runs, and a run in the main session re-sends the conversation, which OpenClaw's own reference puts at about 100K tokens. At the Kunavo input rate for Claude Sonnet 5 that is about $6.72 a day before any real work, and about $0.34 with isolatedSession, which cuts a run to a few thousand tokens. The work on top is mostly cache reads when requests arrive less than five minutes apart.
What happens to OpenClaw when the API balance runs out?
Kunavo refuses the request with HTTP 402 and charges nothing for it. OpenClaw treats a billing failure as a reason to fail over: its documentation says the credential is disabled for ten minutes, the run moves to the next model in agents.defaults.model.fallbacks, and topping up does not clear the window by itself. Two settings on the Kunavo side keep an agent from getting there: auto-recharge charges a saved card when the wallet runs low, so a request that would have been refused is served instead, and a monthly limit on the agent's own key caps what a runaway loop can spend.
Which model should the OpenClaw heartbeat use?
The cheapest one that can read the heartbeat prompt and decide that nothing needs attention. heartbeat.model takes a provider/model reference — for example kunavo/claude-haiku-4-5. Pair it with isolatedSession: true: OpenClaw's heartbeat page warns that a heartbeat which switches a shared session to a smaller model can leave that model in place for the next real turn, and an isolated session avoids that.