On a 128 GB M3 Max running OpenClaw 2026.9.7 with Ollama, gemma4 — the model OpenClaw's own Ollama setup suggests — passed all 12 graded runs across four agent tasks, and gpt-oss:20b passed 11 of 12 while making nearly three times as many failed tool calls. Both took a median of about 54 seconds per task. We ran each model three times on each task in a fresh workspace, graded the result by script, and recorded OpenClaw's own tool and token counts. Measured October 1, 2026.
For what OpenClaw itself costs and what a hosted model would cost instead, see OpenClaw pricing; for running a local model with an API fallback, multiple agents and models in OpenClaw.
What was run
| Machine | Apple M3 Max, 128 GB unified memory |
| OpenClaw | 2026.9.7 from npm, Node 24.21.0, openclaw agent --local --json, one turn per run |
| Ollama | v0.35.0 release binary, native API (no /v1) |
| Models | gemma4:latest — 7.5B, Q4_K_M, 6.6 GB on disk; gpt-oss:20b — 20.9B, MXFP4, 13.8 GB |
| Context | OpenClaw contextWindow 32,768 for both; Ollama loaded both at 131,072 |
| Tool execution | OpenClaw's Docker sandbox, workspace mounted read-write |
| Runs | 4 tasks × 3 repeats × 2 models = 24, each in a fresh workspace with fresh OpenClaw state |
The four tasks
| Task | What the agent had to do | Passes when |
|---|---|---|
| Read | Read a 301-line service log and name the service on its only ERROR line | The reply contains the right service name |
| Fix | Run a unit-test file, fix the bug in the module it tests, rerun — without touching the tests | Tests exit 0 and the test file is byte-identical |
| Count | Count data rows in five CSV files and write the counts to counts.json | counts.json equals the expected object |
| Two-file fix | Two independent bugs in two modules behind three failing tests; fix both without touching the tests | All tests exit 0 and the tests are byte-identical |
Results
| Model | Task | Passed | Median time | Median tool calls | Failed tool calls (3 runs) | Median tokens in / out |
|---|---|---|---|---|---|---|
gemma4:latest | t1-read | 3/3 | 50.2 s | 1 | 0 | 21,438 / 19 |
gemma4:latest | t2-fix | 3/3 | 47.5 s | 4 | 0 | 13,387 / 866 |
gemma4:latest | t3-count | 3/3 | 37.9 s | 7 | 1 | 13,073 / 513 |
gemma4:latest | t4-multi | 3/3 | 75.1 s | 10 | 6 | 16,417 / 2,441 |
gpt-oss:20b | t1-read | 3/3 | 38.7 s | 3 | 5 | 16,971 / 508 |
gpt-oss:20b | t2-fix | 3/3 | 54.2 s | 7 | 6 | 12,409 / 1,193 |
gpt-oss:20b | t3-count | 2/3 | 76.5 s | 10 | 5 | 13,247 / 1,713 |
gpt-oss:20b | t4-multi | 3/3 | 65.1 s | 11 | 3 | 13,065 / 1,380 |
Times are wall-clock per run, including about five to seven seconds of OpenClaw start-up. "Tokens in" is uncached input as OpenClaw recorded it; the context it re-sent from cache is on top of that and is counted in the cost section below.
What the numbers say
- Both models are usable for short agent tasks. 23 of 24 runs passed, including the two-file fix, which needs reading several modules, editing two of them and rerunning tests.
- gemma4 used tools more cleanly. On the read task it made exactly one tool call every time; gpt-oss made three, two of which typically failed before it read the file. Over all twelve runs gpt-oss had 19 failed tool calls to gemma4's 7, and averaged 9.1 assistant turns per task to gemma4's 6.8.
- The one failure was gpt-oss on the count task: after thirteen tool calls it wrote a malformed counts.json and ended the turn without a reply.
- Speed is not the differentiator here. Median 53 s for gemma4 and 54 s for gpt-oss across all runs; the slowest single run was gemma4's 133 s on the two-file fix, with seventeen tool calls.
- The bigger model did not buy accuracy on these tasks. gpt-oss:20b is nearly three times gemma4's parameter count and twice its disk size.
Twelve runs per model is enough to see a pattern on these tasks and not enough to rank models generally. Long tasks, larger codebases and other quantizations were not tested.
The setup that mattered
- Native Ollama URL. OpenClaw's documentation says the
/v1OpenAI-compatible URL "breaks tool calling and models can emit raw tool-call JSON as plain text"; the runs usedbaseUrlwithout/v1andapi: "ollama". - Pin
contextWindowon hand-written entries. In a smoke run, an explicit model entry without it resolved to a 200,000-token context; OpenClaw's own Ollama setup writes 32,768 for local models, which is what these runs used. - Ollama's context is separate. Ollama 0.35.0 loaded both models at 131,072 tokens regardless; that, not OpenClaw's budget, decides memory.
- Sandbox the shell. A local model running shell commands on your machine is the real risk here. With
sandbox.mode: "all", exec ran in OpenClaw's Docker image with only the workspace mounted; the image is built once with thedocker buildcommand in OpenClaw's sandboxing docs.
// ~/.openclaw/openclaw.json — the runs used one model per config; the
// fallbacks line shows the pattern and was not part of the measured runs
{
models: { providers: { ollama: {
apiKey: "ollama-local",
baseUrl: "http://127.0.0.1:11434", // native Ollama URL — no /v1
api: "ollama",
timeoutSeconds: 600,
models: [
{ id: "gemma4:latest", name: "gemma4:latest", contextWindow: 32768, maxTokens: 8192,
params: { keep_alive: "30m" } },
{ id: "gpt-oss:20b", name: "gpt-oss:20b", contextWindow: 32768, maxTokens: 8192,
params: { keep_alive: "30m" } },
],
} } },
agents: { defaults: {
model: { primary: "ollama/gemma4:latest", fallbacks: ["ollama/gpt-oss:20b"] },
sandbox: { mode: "all", workspaceAccess: "rw" }, // shell runs in Docker
} },
tools: { exec: { mode: "full" } },
}What local costs
A local model swaps a per-token bill for time, disk and electricity. Per run, OpenClaw recorded about 16,415 uncached input tokens, 78,469 tokens of re-sent cached context and 1,142 output tokens for gemma4, and 13,761, 88,688 and 1,192 for gpt-oss. As rough arithmetic — tokenizers differ between models — the same counts at Claude Haiku 4.5's Kunavo rates ($0.70 in, $0.07 cached, $3.50 out per million) come to about $0.021 and $0.020 a run. Power was not measured; the electricity per run is watts × seconds ÷ 3,600,000 × your price per kWh.
The practical pattern is a local primary with a hosted fallback for the turns a local model fails. OpenClaw takes a fallbacks list per agent; a Kunavo key works as an OpenAI-compatible provider beside Ollama, billed per token from a prepaid balance. Nobody at Kunavo has run OpenClaw against its endpoint — these runs were entirely local.
FAQ
What is the best local model for OpenClaw?
Of the two we measured, gemma4 — OpenClaw's own suggested Ollama model. On an M3 Max with 128 GB, running OpenClaw 2026.9.7 and Ollama 0.35.0, gemma4 (7.5B, Q4_K_M) passed all 12 graded runs across four tasks; gpt-oss:20b (20.9B, MXFP4) passed 11 of 12. Both took a median of about 54 seconds per task, but gpt-oss made 19 failed tool calls to gemma4's 7. That is a two-model result on short tasks, not a ranking of every local model.
Can OpenClaw run fully offline with Ollama?
Yes for the model calls: with the provider pointed at a local Ollama host, every model request in these runs went to 127.0.0.1. Use the native URL, http://127.0.0.1:11434, not the /v1 OpenAI-compatible one — OpenClaw's Ollama documentation says /v1 breaks tool calling and can make models print raw tool-call JSON as text. Skills or tools that reach the internet still need it, and OpenClaw's Docker sandbox image has to be built once (it is not pulled automatically).
How much memory does a local OpenClaw model need?
On disk, gemma4 is 6.6 GB and gpt-oss:20b 13.8 GB; Ollama reported gpt-oss using 13.7 GB once loaded. Context is the other half: Ollama 0.35.0 loaded both with a 131,072-token context on this machine, while OpenClaw's own budget was the 32,768 we set. On a 128 GB Mac both fit at once with room to spare; on a 16 GB machine gpt-oss:20b alone would be tight. Smaller machines were not tested.
Is a local model free compared with an API?
Not free, just billed differently: you pay in time, disk and electricity instead of per token. Each run here used about 95,000–105,000 tokens counting OpenClaw's re-sent context — on a metered API at Claude Haiku 4.5 rates that would be roughly $0.021 a run, as rough arithmetic across different tokenizers. A local run took about 54 seconds; power draw was not measured, so the electricity side is a formula: watts × seconds ÷ 3,600,000 × your price per kWh.
Why set contextWindow when adding Ollama models to OpenClaw by hand?
Because an explicit model entry without it resolved to a 200,000-token context in our smoke run — far beyond what most local models handle — whereas OpenClaw's own Ollama setup writes 32,768 for local models. Pinning contextWindow (and maxTokens) keeps OpenClaw's compaction budget realistic. It does not change Ollama's own num_ctx: Ollama still loaded the models at 131,072 here.
Run on October 1, 2026 on an Apple M3 Max (128 GB): OpenClaw 2026.9.7 (npm, Node 24.21.0), Ollama v0.35.0 (release binary), gemma4:latest and gpt-oss:20b pulled from the Ollama registry and checked by sha256. Each of 24 runs used a fresh workspace and fresh OpenClaw state, exec in OpenClaw's Docker sandbox, and a script grader; tool calls and token counts are OpenClaw's own JSON output. Task fixtures, graders and the adapter script are published with this page's evidence. Not measured: power draw, long tasks, other quantizations or smaller machines.