DeepSeek Harness costs nothing to run; the bill is the model's, and on its built-in route that is DeepSeek's API — deepseek-flash at $0.30 per million input tokens and $1.20 per million output at peak, half that off-peak, and $0.006 per million on a cache hit. What makes the harness itself matter to the bill is what it sends: in a recorded run of 0.2.0-rc.2, every turn carried 24 tool definitions, asked for up to 256,000 output tokens with thinking on, and every new session added a title request. Prices read October 1, 2026.
For dsh against Anthropic's agent, see DeepSeek Harness vs Claude Code; for pointing it at another endpoint, Kunavo's DeepSeek Harness setup page.
The harness is free; the model is not
DeepSeek Harness (dsh) is MIT-licensed and installs from npm as @deepseek-ai/dsh; 0.2.0-rc.2 (September 29, 2026) is the current tag, and the README still calls it a developer preview with "compatibility-breaking changes" ahead. Out of the box it routes to DeepSeek's official API with the model deepseek-flash, over DeepSeek's Anthropic-format endpoint, with your DeepSeek key. You can add other providers; you cannot buy anything from dsh.
DeepSeek's prices
| Model | Input, cache hit | Input, cache miss | Output |
|---|---|---|---|
deepseek-flash (DeepSeek-V4.1-Flash) | $0.006 peak / $0.003 off-peak | $0.3 / $0.15 | $1.2 / $0.60 |
deepseek-v4-pro | $0.044 peak / $0.022 off-peak | $1.32 / $0.66 | $3.96 / $1.98 |
Per million tokens, from DeepSeek's pricing page on October 1, 2026. Off-peak rates are half of peak; peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, excluding Chinese public holidays. deepseek-flash is served by DeepSeek-V4.1-Flash; the older deepseek-v4-flash name is still accepted and billed at the Flash price. Charges come off a prepaid DeepSeek balance.
What one turn sends — measured
We ran dsh 0.2.0-rc.2 headless against a local stand-in endpoint that records requests, with one task that reads a file and answers — three fresh sessions per route, all nine completing the tool round trip, request sizes identical run to run. On the built-in DeepSeek route:
| Request | Size | What is in it |
|---|---|---|
| Agent turn, first call | 65,101 bytes | 24 tool definitions (19,304), system prompt (2,744), task plus runtime context, max_tokens: 256000, thinking enabled, output_config.effort: "high", and two non-model fields: dsh_session_log (34,168) and dsh_plugin_packages (6,150) |
| Session title | 39,778 bytes | max_tokens: 64, thinking disabled, the first prompt — and the same two non-model fields |
| Agent turn after the tool result | 36,097 bytes | The conversation so far, with the session log now sent incrementally |
Three things follow for cost. The title call is per session, small in tokens but real, and it goes to the same model. Thinking is on at high effort by default, and reasoning tokens bill as output, so the effort setting in the model picker is a cost control. The prompt prefix is stable — tools and system prompt were byte-identical across runs — which is what DeepSeek's cache-hit price rewards. The two dsh_ fields are not model input; whether DeepSeek counts them toward anything is not documented, and nothing here assumes it does. These are request bytes from a stand-in, not DeepSeek token counts.
What a session costs
Token arithmetic, not a measured bill. One session that sends 200,000 input tokens and receives 12,000 output tokens, peak prices:
| Route | No cache | 80% of input cached | Billed by |
|---|---|---|---|
deepseek-flash, peak | $0.074 | $0.027 | DeepSeek |
deepseek-v4-pro, peak | $0.312 | $0.107 | DeepSeek |
deepseek-flash, off-peak | $0.037 | $0.014 | DeepSeek |
| Claude Haiku 4.5 through Kunavo | $0.182 | — | Kunavo |
| Claude Sonnet 5 through Kunavo | $0.364 | — | Kunavo |
The cached column assumes DeepSeek's automatic cache serves 80% of input, which a stable prefix makes plausible but which only your own usage page can confirm. The Kunavo rows are live catalog rates without caching, for readers who want Claude behind the same harness; Kunavo does not sell DeepSeek models, bills per token from a prepaid balance with a $10 minimum top-up (billing), and treats its catalog amount as a floor rather than a cap. The price gap is real — and so is the gap between cheapest rate and cheapest finished task.
Other providers, and what changes
A custom model API takes a base URL and one protocol. In the run, openai-completions with a base URL ending in /v1 posted to /v1/chat/completions; anthropic-messages with the bare root posted to /v1/messages?beta=true, and with /v1 on the end posted to /v1/v1/messages — a 404 on any real gateway. These routes asked for max_tokens: 32768 rather than 256,000, sent neither dsh_ field, and their title request was under 800 bytes.
The session-log field is the privacy half of the same choice. On the built-in route the dsh-session-log-deepseek plugin is on by default and sends session events, including your working-directory path, with each request — to whatever base URL the route uses. Turn it off under Settings → General, "Upload Session Log when using the official model API". dsh's OpenTelemetry feedback export is a separate default, to DeepSeek's collector; DSH_TELEMETRY_MODE=DISABLED turns it off, which is how these runs were made.
FAQ
How much does DeepSeek Harness cost?
The harness is free: DeepSeek Harness (dsh) is MIT-licensed, and 0.2.0-rc.2 on npm (September 29, 2026) is still a developer preview. What you pay for is the model. Its built-in route sends DeepSeek's own API, where deepseek-flash lists at $0.30 per million input tokens on a cache miss, $0.006 on a cache hit and $1.20 per million output tokens at peak, and exactly half that off-peak; deepseek-v4-pro is $1.32, $0.044 and $3.96 at peak. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday. There is no dsh subscription.
What does one DeepSeek Harness turn send?
On the built-in DeepSeek route in 0.2.0-rc.2, the first request of a one-line task was 65,101 bytes: 24 tool definitions (19,304 bytes), a 2,744-byte system prompt, the task with a runtime-context note, and two fields that are not model input — dsh_session_log (34,168 bytes on that first call) and dsh_plugin_packages (6,150). It asks for max_tokens 256,000 with thinking on and effort "high". Each new session also sends a short title request. Those figures are request bytes recorded by a stand-in endpoint, not DeepSeek's token count.
Does DeepSeek Harness upload my session?
On its built-in DeepSeek route, yes by default. The dsh-session-log-deepseek plugin is enabled out of the box and adds a dsh_session_log field — the session's events, including the working-directory path — to normal agent, compaction and title requests, up to 8 MiB per request; the model does not see it. In the run here the field went to whatever DEEPSEEK_BASE_URL pointed at, not only to DeepSeek. The Web UI's Settings → General → "Upload Session Log when using the official model API" switch turns it off, as does enabled: false for that plugin. Custom pi-ai providers did not carry it.
What is the cheapest way to run DeepSeek Harness?
On list price, deepseek-flash off-peak with a warm cache: $0.15 per million input tokens on a miss, $0.003 on a hit and $0.60 per million output. For a session that sends 200,000 input tokens and receives 12,000 output, that is about $0.014 off-peak with 80% of the input cached, $0.037 off-peak with no cache hits, and $0.074 at peak with none. Cheapest listed rate and cheapest way to finish the task are different questions: a model that needs three attempts costs more than one that needs one.
Can DeepSeek Harness use a different API, like Claude through a gateway?
Yes. A custom model API in Settings → Models (stored in the profile's cordis.patch.yml) takes a base URL and one protocol: openai-completions, openai-responses or anthropic-messages. In the 0.2.0-rc.2 run, openai-completions with a base URL ending in /v1 posted to /v1/chat/completions, and anthropic-messages posted to /v1/messages?beta=true when given the bare root — given a base URL ending in /v1 it posted to /v1/v1/messages, which a real gateway would 404. Kunavo's DeepSeek Harness page has the openai-completions setup.
Run on October 1, 2026: @deepseek-ai/dsh 0.2.0-rc.2 from npm, dsh --profile headless, telemetry disabled, fake keys, against a local recording stand-in that returns one tool call and a fixed answer — three routes, three sessions each. The stand-in is not DeepSeek, Kunavo or a model: it shows what dsh sends, not what a provider bills or whether it accepts the extra fields. Prices from DeepSeek's pricing page and Kunavo's live catalog the same day; package behaviour from the shipped READMEs of dsh-llm-deepseek and dsh-session-log-deepseek.