"Context compression timed out before it could commit" means Hermes's summary model — the one under auxiliary.compression, not your main model — made no progress before Hermes's compression budget ran out, and the turn was stopped without the main call being sent. It is version-dependent: we reproduced both of its messages on Hermes Agent v0.21.3 and recovered from them by making the summariser answer and retrying in the same session, while on v0.21.5 the same stall produced a long wait instead of the error. Tested October 1, 2026 against a recording stand-in for both models, not a real provider.
If you are setting Hermes up against your own endpoint rather than debugging it, start with Hermes Agent with a custom API.
The two messages, and what each one tells you
| What you see | When it appeared | Main call sent? |
|---|---|---|
| "Context compression timed out before it could commit while the request was still approximately 83,485 tokens. The provider call was not sent. Run /compress and wait for it to finish, then retry." | v0.21.3, summariser silent, request (~83,485 tokens) larger than the model's 64,000-token window | No |
| "Context compression timed out without reducing this conversation. No messages were dropped. Start a fresh session with /new, or check auxiliary.compression before retrying /compress." | v0.21.3, summariser silent, request (~57,489 tokens) past the 54,400-token compression trigger but inside the window | No |
The token figure in the first message is Hermes's own estimate of the request it would have sent. Both turns ended in the log as reason=context_compression_timeout with api_calls=0, after a warning of the form "Context compression made no progress for 5.0s" — five seconds because the test set compression.context_timeout_seconds: 5 to keep runs short; the default is 120.
Which version you are on decides what happens
| Summariser state | v0.21.3 (September 14) | v0.21.5 (September 24) |
|---|---|---|
| Silent, request inside the window | Turn stopped — second message above | Waited for it (30 s and 90 s holds), compressed 9 messages to 5, sent the request |
| Silent, request over the window | Turn stopped — first message above | Not run with a silent summariser; the docs bound this case by one inactivity budget and a deterministic fallback summary |
| Answering errors (404) | Two attempts, compression skipped, request sent (tested inside the window) | Same; an over-window request was sent too, which a real provider would reject |
| Responsive | Compressed and sent | Compressed and sent |
The boundary is in the source. In v0.21.4 (tag v2026.9.21) the function that stops a turn after a timed-out compression was narrowed to requests above the model's window, citing issues #113646 and #114594; v0.21.3 stops every one. v0.21.5's configuration documentation then adds that the compression wait "is floored at the auxiliary compression request's own timeout (auxiliary.compression.timeout, minimum 300s)" — which is why our five-second setting did nothing there. So on a current build, a slow summariser costs you minutes of waiting rather than an error, and a dead one is skipped.
Shortest diagnosis
- Check the version with
hermes --version. On v0.21.3 or earlier, both messages above are expected behaviour for a stalled summariser. - Find the summariser in
~/.hermes/logs/agent.log: a line "Auxiliary compression: using <provider> (<model>) at <url>" names the model and endpoint that timed out. With noauxiliary.compressionblock it inherits your main model. - Read the trigger line, "Pre-API compression: ~N request tokens >= T threshold (context=W)". If N is above W, the turn cannot go out unshrunk on any version.
- Check the summariser's window. Hermes's documentation requires it to be at least as large as the main model's, because it is sent the whole middle of the conversation.
Recovery that was verified
In every v0.21.3 case above, restarting the stand-in with a summariser that answered and sending one more turn in the same session (--resume) produced a normal reply, with the conversation intact. In practice that means one of: point auxiliary.compression at a fast model you can reach, fix the endpoint it points at, or raise compression.context_timeout_seconds if your summariser is slow but healthy — a local model, for instance. Then retry in the same session.
# ~/.hermes/config.yaml — the summariser is its own model, with its own budget
auxiliary:
compression:
base_url: https://api.kunavo.com/v1 # overrides provider; any OpenAI-compatible endpoint
api_key: sk-kn-...
model: claude-haiku-4-5 # fast, and a context window >= your main model's
timeout: 300
compression:
context_timeout_seconds: 120 # inactivity budget for the summary (default)
context_total_ceiling_seconds: 600Two things were not tested: Hermes's Desktop and messaging-gateway surfaces, which have their own hygiene compression with different budgets, and a real provider rejecting an over-window request. The runs were one-shot hermes chat -Q turns.
What it costs
A stopped turn sends no main-model request, so the main model bills nothing for it. The summary request was sent — about 19,000 characters of prompt in these runs — and a provider bills it if it eventually completes. The retry then pays for one summary and one main call. A small, fast summariser keeps both the wait and that overhead down; on Kunavo, Claude Haiku 4.5 lists at $0.70 per million input tokens and $3.50 per million output, billed per token from a prepaid balance. Nobody at Kunavo has run Hermes against its endpoint; the runs here used a local stand-in.
FAQ
What does "Context compression timed out before it could commit" mean in Hermes?
Hermes tried to shrink the conversation before sending your turn, the summary model it uses for that made no progress within its time budget, and the request was too large to send unshrunk — so Hermes stopped the turn without calling your main model at all. Reproduced on Hermes Agent v0.21.3 with a stalled summariser and an 83,485-token request against a 64,000-token window; the log line for the turn read api_calls=0. Your main model and its API timeout are not the problem. The summariser is: the model under auxiliary.compression in config.yaml.
How do I fix a Hermes context compression timeout?
Make the summariser answer, then retry in the same session. In the reproduction, pointing auxiliary.compression at a responsive model and sending the next turn in the same session worked on v0.21.3 every time, with the history intact — the message itself says no messages were dropped. Upgrading also changes the picture: from v0.21.4 a timed-out compression no longer stops a request that still fits the model's window, and in v0.21.5 the summary wait is floored at the summariser request's own timeout, at least 300 seconds. Starting a fresh session with /new works too, at the cost of the conversation's context.
Is the Hermes compression timeout fixed?
Partly, and it changed shape. Up to v0.21.3 (September 14, 2026) any timed-out preflight compression stopped the turn. v0.21.4 (September 21) stops only a request that exceeds the model's context window. In v0.21.5 (September 24) a stalled summariser held for 30 and then 90 seconds did not stop the turn at all: Hermes waited, compressed and sent the request — so on a current version the symptom of a slow summariser is a long pause, not this error. A summariser that is dead rather than slow is different: it is retried, then skipped, and the turn goes ahead uncompressed.
Does raising HERMES_API_TIMEOUT help?
No. That variable governs the main model call, 1,800 seconds by default. Compression has its own budgets in config.yaml: compression.context_timeout_seconds (an inactivity budget, default 120 seconds), compression.context_total_ceiling_seconds (default 600) and auxiliary.compression.timeout for the summary request itself. Hermes's documentation adds a requirement that matters as much as speed: the summary model's context window must be at least as large as the main model's, because it receives the whole middle of the conversation.
Did the timed-out turn cost anything?
Not on the main model: the turn ended before the main call was sent. The summary request was sent, though, and a summariser that eventually finishes is billed by its provider like any other call — in the runs here the summary prompt was about 19,000 characters. The retry then pays for one summary and one main call. That is why pointing compression at a small, fast model is cheaper as well as quicker: on Kunavo, Claude Haiku 4.5 lists at $0.70 per million input tokens.
Reproduced on October 1, 2026 with Hermes Agent v0.21.3 (tag v2026.9.14) and v0.21.5 (tag v2026.9.24), installed from their release sources, against a local recording stand-in serving both the main and the summary model (held responses for a stalled summariser, 404 for a failing one), with a 64,000-token window and four resumed turns of filler before the test turn. The version boundary was read from agent/turn_context.py at tags v2026.9.14, v2026.9.21 and v2026.9.24, and the budgets from each tag's configuration documentation. The stand-in is not a model or a provider: it accepts any request size, so provider-side context errors were not reproduced.