Docs
Hermes Agent
Hermes Agent takes any endpoint as a custom provider, through hermes model or a few lines of config.yaml. For an agent that runs on its own, the lines are the short part: this page also covers which side places the cache breakpoints on each wire, what scheduled jobs and side tasks add to a day, and what a 402 does to a turn.
A named provider in ~/.hermes/config.yaml — api https://api.kunavo.com, transport anthropic_messages, selected with provider: custom:kunavo — puts Hermes Agent on Claude over the Messages wire, where it sends its own cache markers and output limit.
# ~/.hermes/config.yaml
providers:
kunavo:
api: https://api.kunavo.com # origin — the Anthropic SDK adds /v1/messages
key_env: KUNAVO_API_KEY # the variable's NAME; the key goes in ~/.hermes/.env
transport: anthropic_messages
models:
claude-sonnet-5:
context_length: 1000000
prompt_caching: true
claude-haiku-4-5:
context_length: 200000
prompt_caching: true
model:
default: claude-sonnet-5
provider: custom:kunavoapi is the origin — https://api.kunavo.com, no /v1. The Anthropic SDK that Hermes uses for this transport appends /v1/messages itself, and Hermes' documentation says Hermes removes a trailing /v1 before handing the URL to that SDK, so the origin is the form that is right either way. The OpenAI-compatible wire, further down, is the one that keeps the suffix.transport: anthropic_messages is worth writing by hand. Hermes can detect the wire from the URL, but the only rule its documentation spells out is a path ending in /anthropic, which this base URL does not have.prompt_caching: true makes the cache markers explicit for that model on this entry, and context_length is the catalog's window — 1,000,000 tokens on Claude Sonnet 5, at a flat rate. Hermes compresses at half the window by default, which on a window this size is late; the cost section below shows the setting that brings it forward.sk-kn-) and add credit from $10 — calls are paid from that balance, and failed calls are not billed. The dashboard then opens on the Hermes Agent setup.Step by step
- Create a key at
/app/keysand copy it — it is shown once. - Store it where Hermes keeps secrets:
hermes config set KUNAVO_API_KEY sk-kn-...writes it to~/.hermes/.env. Thekey_envline in the block names that variable; the key itself never goes intoconfig.yaml. - Add the block to
~/.hermes/config.yaml—hermes config editopens it. If amodel:section is already there, replace itsdefaultandproviderand leave the rest. - Or let the wizard write it: run
hermes modelin a terminal, outside any chat session, choose Custom endpoint (self-hosted / VLLM / etc.), and answer its prompts — the API base URL, the key, the model name, the API mode and the context length. - Start
hermesand read the banner: the model and its context window are shown there, and both should match the block. - Send two messages and open
/usageto see what the turns used. To change model inside a session, use/model custom:kunavo:claude-opus-5-5.
Checked against Hermes Agent's AI providers page on October 5, 2026. Third-party settings move; if a field name here no longer matches what you see, that page is the authority, not this one.
Verify before you debug the client
One request settles whether a failure is the endpoint, the key, or the configuration file. If this returns JSON, the same base URL and key work in Hermes Agent.
# Settles whether a failure is the endpoint, the key, or the client.
curl -sS https://api.kunavo.com/v1/messages \
-H "Authorization: Bearer sk-kn-..." \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model":"claude-sonnet-5","max_tokens":16,"messages":[{"role":"user","content":"ping"}]}'Which model id to put in the field
Every text model is reachable as a model id — the live list is GET /v1/models, and the catalog with prices is on the models page. Rates are USD per 1M tokens, input / output.
| Model id | Kunavo in / out | Where it fits in Hermes Agent |
|---|---|---|
claude-sonnet-5 | $1.40 / $7.00 | the main model — conversations, tool loops and delegated work |
claude-opus-5-5 | $2.80 / $14.00 | the step up for long or difficult tasks; add it under models and switch with /model |
claude-haiku-4-5 | $0.70 / $3.50 | auxiliary tasks and scheduled jobs — compression, titles, cron.model |
claude-fable-5 | $7.00 / $35.00 | the top tier — price a day on it with the table below before leaving an agent on it |
The OpenAI-compatible wire
The same key reaches every other model family through /v1/chat/completions, and for that Hermes' shortest documented form is enough — a model: block with provider: custom and a base_url, which is also what hermes model asks for:
# ~/.hermes/config.yaml — the bare form, for an OpenAI-compatible endpoint
model:
default: gpt-6-sol
provider: custom
base_url: https://api.kunavo.com/v1 # this wire keeps /v1
key_env: KUNAVO_API_KEY
context_length: 1050000The base URL keeps /v1 here, the form Hermes' documentation uses for its local-server examples, and context_length pins the window so Hermes does not have to detect it. To keep both wires configured at once, give this one its own named entry — transport: chat_completions — and switch with /model custom:<name>:<model>.
/v1/chat/completions Kunavo gives a Claude request that names no limit 4,096 output tokens, which cuts a long reply or a large tool call short. On the Anthropic wire Hermes supplies max_tokens itself. If you do run Claude here, a named entry's extra_body is the documented way to add a field such as max_tokens to every chat-completions request. Thinking controls are not forwarded for Claude on this wire either.Prompt caching on each wire
On the Anthropic wire Hermes attaches the cache markers itself. For a custom provider its configuring-models page documents the switch used in the block above — prompt_caching: true on the model — and says the layout follows the transport: native blocks on anthropic_messages, the envelope layout on the OpenAI-compatible wire. Kunavo's /v1/messages forwards the body as it was sent and adds no breakpoint of its own, so on this wire the markers are Hermes' or there are none.
The lifetime is the part Hermes' documentation leaves open for a custom endpoint. prompt_caching.cache_ttl — 5m, 1h or auto — is described for Claude through the native Anthropic API, OpenRouter and Nous Portal, with nothing said about other endpoints. Kunavo forwards whichever marker arrives and bills the write at the same rate, so read the answer off your own usage: a turn after a ten-minute pause that still shows a cache read was holding a one-hour entry.
On the OpenAI-compatible wire Kunavo places the breakpoints for Claude models itself — on the system prompt, the tool definitions and the end of the conversation — once the prompt is long enough to cache, whether or not the client sends a marker. GPT models are cached implicitly by their vendor.
Whichever side places the breakpoints, the bill reads the same. On Claude Sonnet 5 a cache read is $0.14 per 1M tokens against $1.40 for fresh input, and a cache write is $1.75 — the write premium Claude carries over input, charged at that same rate when the entry asks for a one-hour lifetime. An entry lasts five minutes and every read renews it, so what an agent pays depends less on the model than on whether its next request arrives inside that window. Every model's cache rates are on the prompt caching page.
One Hermes behaviour costs more than any rate: its documentation notes that a model switch in the middle of a session, an automatic fallback or a credential rotation resets the prompt cache, so the next message re-reads the whole conversation at the full input price. Pick the model before a long session starts.
What an always-on agent costs per day
On Hermes, what bills while nobody is typing is whatever you scheduled, plus the side tasks every conversation sets off. Its cron documentation says each scheduled run starts a fresh session, so a run bills its whole prompt — instructions, tool schemas, attached skills — every time it fires. The size of that prompt depends on your setup, so the table is an assumption stated outright: 20,000 tokens a run and one run every 30 minutes, 48 a day. Replace both with your own figures.
| Model on the job | Input rate per 1M tokens | 48 runs a day |
|---|---|---|
claude-haiku-4-5 | $0.70 | $0.67 |
claude-sonnet-5 | $1.40 | $1.34 |
claude-opus-5-5 | $2.80 | $2.69 |
claude-fable-5 | $7.00 | $6.72 |
Three settings move that figure, all from Hermes' own documentation. A scheduled job runs on its per-job model, else on cron.model, else on the main model — so hermes config set cron.model claude-haiku-4-5 takes every unpinned job off the expensive tier. A job script that prints {"wakeAgent": false} skips the model for that tick, and a no-agent job never calls one. And the side tasks — compression, titles, vision — run on the main model unless auxiliary routes them elsewhere:
# ~/.hermes/config.yaml — what decides the cost of an unattended day
compression:
threshold_tokens: 256000 # compact here, not at half of a 1M window
auxiliary:
compression:
provider: kunavo # the named entry above
model: claude-haiku-4-5 # summaries on the cheapest tier
title_generation:
provider: kunavo
model: claude-haiku-4-5threshold_tokens is the one that matters on a large window. Compaction starts at half the context length by default, and Hermes' documentation gives this setting as the way to put a fixed ceiling on what one call can cost.
The hours when the agent is actually working are the other half of the bill, and there the cache decides it. Take 100 consecutive requests, each re-sending a 100,000-token context with 2,000 new tokens on top and returning 800 tokens of output. On Claude Sonnet 5 that is about $2.31 while the context is read from cache, and about $14.84 when every request bills it as fresh input. Same work, same model: the difference is whether the breakpoints are there and the requests are less than five minutes apart.
For scale, measured and not assumed: across the Kunavo accounts that run an always-on agent, a median active day has cost $12.67 and a 90th-percentile day about $163. Those are amounts billed up to October 5, 2026, at the rates in force on each day. It is a small group, so read it as the width of the range, not as a forecast for your agent.
When the balance runs out
Kunavo is prepaid: every call is paid from the wallet, and an agent that works while you sleep empties it while you sleep. A request the wallet cannot cover is refused with HTTP 402 and the code insufficient_balance, on either wire, and nothing is charged for it. The refusal comes before the wallet reads zero: each request first reserves its worst-case cost, its prompt plus the largest reply it is allowed to produce, so the larger the output cap an agent asks for, the earlier its calls start to bounce. The error says how far short it was, in balance_usd and needed_usd.
Hermes Agent's answer to a failing provider is a fallback chain: fallback_providers in config.yaml, managed with hermes fallback and tried turn by turn. Its documentation lists rate limits, server errors, authentication failures and 404s as what triggers it for the main model, and names HTTP 402 among the capacity errors that move a side task down its chain; it does not say what a turn does with a 402 when no fallback is configured. Plan for the plain reading — the turn fails, and a scheduled job with it — and remember that a turn which does fall back starts with a cold prompt cache.
Two settings keep an unattended agent out of that state, and they do different jobs:
- Auto-recharge, under Billing. Save a card once and set three numbers: the balance below which to top up, the amount to add each time, and a monthly cap. The wallet then refills within seconds of a call that takes it under the threshold. A request that arrives while the wallet is still short waits for that charge and is then served instead of refused. A
402still comes back when the charge cannot be made — a declined card, the monthly cap reached — or when one request reserves more than the wallet holds after the top-up. It needs a card or Link — Alipay, WeChat Pay, Pix and the other local methods cannot be charged automatically. - A monthly limit on the key, under API Keys. Give the agent a key of its own and set the most that key may spend in a calendar month. Past that figure its calls are refused with a
402and nothing is charged, while your other keys keep working. That is the ceiling a runaway loop needs, and one the wallet cannot provide, because every key draws on the same wallet.
Set the recharge threshold above what one request reserves, and size the amount from a day of your agent, not from the minimum: the smallest top-up is $10, and the median always-on day above is $12.67. The limits on auto-recharge are on the billing page, and the whole error body is on the errors page.
Related guides
- Hermes Agent custom API — why the outbound provider is not the inbound API server, the transport field explained, and the checks to run on the first call.
- Hermes Agent pricing — what running it costs beyond the token rates.
- Hermes context compression timed out — what the error means when the summariser stalls, and the recovery.
- Hermes vs OpenClaw — and the same setup for the other agent, on the OpenClaw page.
Frequently asked questions
How do I add a custom endpoint to Hermes Agent?
Run hermes model from a terminal, outside any chat session, and choose "Custom endpoint (self-hosted / VLLM / etc.)": it asks for the API base URL, the API key and the model name, then for the API mode and the context length, and saves the result to ~/.hermes/config.yaml. You can also write it by hand — either a model: section on its own, with provider: custom and base_url, or a named entry under providers: with api, key_env and transport, selected with provider: custom:<name>. The /model command inside a session only switches between providers that already exist.
Does the Hermes Agent base URL need /v1?
It depends on the transport. For an OpenAI-compatible endpoint (transport chat_completions) the base URL keeps the suffix, the form Hermes' providers page uses for its local-server examples — for Kunavo, https://api.kunavo.com/v1. For an Anthropic-compatible endpoint (transport anthropic_messages) write the origin — https://api.kunavo.com — because the Anthropic SDK appends /v1/messages itself. Hermes' Microsoft Foundry guide says Hermes strips a trailing /v1 before handing the URL to that SDK, so the origin is the form that is right either way.
Does prompt caching work in Hermes Agent through a custom endpoint?
Yes. Hermes documents a per-model prompt_caching: true setting for custom provider entries, and says the marker layout follows the configured transport: native blocks on anthropic_messages, the envelope layout on the OpenAI-compatible wire. Declaring it for each Claude id makes the behaviour explicit instead of leaving it to detection. On Kunavo's OpenAI-compatible endpoint the gateway also places breakpoints for Claude models by itself, so a chat-completions setup caches even when the client sends no marker.
What does context_length do in Hermes Agent?
It is the total context window Hermes assumes for the model — input and output together — and Hermes uses it to decide when to compress history. Set under model: it is a pin that wins over everything Hermes would otherwise detect; set under providers.<name>.models.<id> it applies to that model on that provider. For a model with a very large window, the cost lever is a separate setting: compression.threshold_tokens makes compaction start at an absolute token count instead of at half the window.
How much does it cost to run Hermes Agent all day?
Count three things. Scheduled jobs: each cron run starts a fresh session and bills its whole prompt — 48 runs a day at an assumed 20,000 tokens each is about $1.34 a day on Claude Sonnet 5 at Kunavo's input rate, and about $0.67 on Claude Haiku 4.5. Side tasks: compression, titles and vision run on the main model unless the auxiliary settings route them elsewhere. And the conversation itself, which is mostly cache reads while turns arrive less than five minutes apart and a full re-read after a longer pause, a model switch or a fallback.
What happens to Hermes Agent when the API balance runs out?
Kunavo refuses the request with HTTP 402 and charges nothing for it. Hermes' answer to a failing provider is its fallback chain — fallback_providers in config.yaml, tried turn by turn — and a turn that falls back starts with a cold prompt cache on the other provider. With no fallback configured, plan for the turn, or the scheduled job, simply failing. Two settings on the Kunavo side keep an agent from getting there: auto-recharge charges a saved card when the wallet runs low, so a request that would have been refused is served instead, and a monthly limit on the agent's own key caps what a runaway loop can spend.