Back to guides
Coding agents·October 1, 2026·9 min read

Warp AI pricing: credits, plans and what your own inference removes

Warp's terminal is free and its AI is metered in credits across three buckets. Your own key or endpoint removes AI credits for the calls it carries — but not Auto, not cloud agents, and on Business not the per-hour platform charge.

Last reviewed on .

Warp's terminal costs nothing; its AI is metered in credits, and the credits are the part worth understanding. As of October 1, 2026, Warp's pricing page gives Build 1,500 credits a month for $20, Max 18,000 for $200 and Business 1,500 per seat for $50, and sells more in Reload packs from 400 credits for $10. Plan credits reset every 30 days; bought credits last 12 months. And a model call you route through your own key or endpoint uses no AI credits at all — but Auto, cloud agents and, on Business, a per-hour platform charge still do.

This page is about that accounting: what each bucket meters, what a credit costs by channel, and exactly which usage your own inference removes. For Warp as a product against Claude Code — licensing, protocols, where the agent harness runs — see Warp vs Claude Code. Everything below was read from Warp's pricing page and documentation on October 1, 2026; nobody here holds a Warp subscription, so no figure is a measured bill.

Which Warp this is

Warp here is the agentic terminal from warp.dev — the desktop app and the Warp Agent CLI. It is not Cloudflare's WARP, the 1.1.1.1 VPN client with its own paid tier, and not the Rust web framework called warp. None of the prices below apply to either.

The plans, and what each includes

PlanMonthlyBilled annuallyIncluded creditsYour own inference
Free$0$0None for the Warp AgentOwn API keys, SuperGrok or a custom endpoint — for individuals and organizations of 10 or fewer employees
Build$20$18 a month1,500 a month — the card says "$20 of included agent usage at API rates"Same as Free
Max$200$180 a month18,000 a month — valued by Warp at $240Same as Free
Business$50 per seat$45 per seat1,500 per seat; up to 25 seatsUnlocked for organizations above 10 employees
EnterpriseCustomCustomCustom shared poolsAdds BYOLLM (inference in your own cloud) and team-managed keys and endpoints

"Pay as you go, starting at" is Warp's own wording on every paid card: the subscription is a floor, and anything past the included credits is bought as Reload credits.

Three buckets, one balance

Warp's documentation meters agent work in three kinds of credit that all draw from the same balance — plan credits first, then Reload credits.

BucketWhat it pays forWhen it is consumed
AI creditsThe model callWhen Warp pays the provider — Warp-managed models, and always for Auto
Compute creditsThe sandbox an agent runs inCloud agent runs on Warp-hosted workers; local runs use your machine and consume none
Platform creditsRun lifecycle, integrations, dashboard, APIs, observability — billed by the agent hourEvery cloud agent run on any plan, and local runs on Business or Enterprise that use your own key, endpoint or BYOLLM

Two details decide real bills. Platform credits accrue per agent hour, partial hours proportionally, and for the Warp Agent only while it is generating — but for Claude Code and Codex run as cloud harnesses, the entire run counts, including time spent waiting on the model provider. And the per-hour rate is not printed where the docs point: the platform-credit page says to see the pricing page for current rates, and on October 1, 2026 the pricing page did not list one.

What a credit costs, by channel

How you get the creditCreditsPricePer credit
Build plan1,500$20$0.0133
Max plan18,000$200$0.0111
Reload pack, paid plan400$10$0.0250
Reload pack, paid plan1,000$20$0.0200
Reload pack, paid plan3,000$50$0.0167
Reload pack, paid plan6,500$100$0.0154
Reload pack, Free plan40020% above the paid-plan rate, per Warpabout $0.030

The per-credit column is division of Warp's list prices, and the Free-plan row applies Warp's stated 20% to the paid rate rather than quoting a printed price. Warp's own tooltip calls Max "about 17% cheaper per credit than buying Reload credits"; plain division against the best Reload tier gives a wider gap, so treat 17% as Warp's framing, not a figure this page derived. What none of these rows tells you is how many credits your work takes: Warp says two similar prompts can use different numbers, that bigger models, more tool calls and more context all raise it, and that there is no formula.

Resets, rollover and the team pool

  • Plan credits reset every 30 days from your subscription or renewal date — not on the calendar month — and do not carry over.
  • Reload credits roll over and stay valid for 12 months from purchase, including after a downgrade to Free. They are refundable only if none has been used.
  • On a team, Reload credits are one shared pool. Any member can buy them; each member spends their own plan credits first, then the shared pool — so one heavy user can spend credits someone else bought. Balances left in personal pools from the May–August 2026 window are used only after the shared pool is empty.
  • Auto-reload is off by default for new subscribers; switched on, it buys your chosen pack whenever the balance drops below 100 credits, under a monthly spend cap that starts at $200 and resets on the calendar month, not the billing cycle.
  • Running out on a paid plan without Reload disables premium models until your next reset; Warp also documents a separate token limit, reported as QuotaLimit, that can stop all models even when credits remain.

When your own inference replaces credits — and when it does not

Warp documents three ways to bring your own model access: an API key for OpenAI, Anthropic or Google (BYOK), a SuperGrok or X Premium subscription for Grok, and a custom inference endpoint — any service that implements OpenAI Chat Completions. For requests routed through any of them, Warp consumes no AI credits and the provider bills you directly. The exceptions are what make a budget honest:

  • Auto always uses Warp credits. To use your endpoint you have to pick the endpoint-routed model in the picker; custom routers cannot resolve to one either.
  • Cloud agents always use Warp credits. A self-serve custom endpoint is stored on your device and never reaches cloud runs; on Enterprise, admins can configure team-managed endpoints that do.
  • Business and Enterprise pay platform credits on local agent runs that use your own key or endpoint. On Free, Build and Max, local runs with your own inference use none.
  • The 10-employee line. Self-serve BYOK and custom endpoints are for individuals and organizations of 10 or fewer employees under Warp's terms; above that, Business or Enterprise is required.
  • The app, not the CLI. Warp's pricing FAQ says custom inference endpoints are available only in the Warp app; the Warp Agent CLI takes API keys and SuperGrok through /api-keys.
  • A public URL, through Warp's servers. The agent harness runs on Warp's backend and calls your endpoint from there, so localhost and private addresses are rejected, and Warp states it cannot enforce Zero Data Retention for that traffic.

Three budgets, worked

These are illustrative, not measured. The token side assumes one developer month of 20M uncached input and 0.8M output tokens through a custom endpoint, at Kunavo's live catalog rates per million tokens; it excludes caching discounts, retries and anything that still runs on Warp credits.

Model through the endpointInput / output per 1MAssumed month of tokens
Claude Haiku 4.5$0.70 / $3.50$16.80
GPT-5.6 Terra$0.70 / $4.20$17.36
Claude Sonnet 5$1.40 / $7.00$33.60
  • A solo developer on Free with an endpoint. Warp charges $0; the model bill is the endpoint's. Auto and any cloud agent run need credits, which on Free means Reload at the higher rate.
  • The same developer on Build. $20 a month buys 1,500 credits for Auto, cloud runs and the models you do not route through your endpoint; the endpoint bill is unchanged. Build is worth it for what the credits and the plan add, not for the endpoint, which Free already allows.
  • A 12-person company. Above the 10-employee line, own inference requires Business: $50 a seat, plus platform credits for every local agent hour that uses the endpoint, at a rate Warp does not print on its pricing page — ask before you commit a team to it.

Setting the endpoint, if you go that way

Warp's custom-endpoint documentation gives the steps: open Settings, search for inference endpoint, add the base URL that exposes /v1/chat/completions and a key, list the model identifiers, and save. Its own local-model example uses a URL ending in /v1.

Warp custom inference endpoint — fields from Warp's docs, not tested here
Warp app  →  Settings  →  search "inference endpoint"

Endpoint URL   https://api.kunavo.com/v1     (the base that exposes /v1/chat/completions)
API key        sk-kn-...                     (stored in your OS keychain, sent per request)
Model IDs      claude-sonnet-5, claude-haiku-4-5  (one per model you want in the picker)

Then pick the endpoint-routed model in the model picker — not Auto.

Kunavo's endpoint is public and OpenAI-compatible, which is what Warp asks for — but that is a match between two sets of documentation. Nobody at Kunavo has run Warp against it, so send one bounded task, confirm it streams and that a tool call returns, and compare Warp's credit chip (which should read zero AI credits for that turn) with the charge in your Kunavo usage log. Kunavo bills per token from a prepaid balance with a $10 minimum top-up and no subscription — see billing.

Measuring your own credits per task

Because Warp publishes no credits-to-tokens ratio, the only reliable number is yours. In the app, hover the credit chip under an agent response for that turn's credits; in the Warp Agent CLI, /cost prints credits per response; Settings > Billing and usage has the running total, with platform, AI and compute credits shown separately. Run the same small task once on a Warp-managed model and once through your endpoint, and you have the one comparison no pricing page can give you.

Related reading: Warp vs Claude Code for the product decision, OpenAI-compatible API for what an endpoint must implement, and the AI agent API directory for other agents that take a custom endpoint without a plan gate.

FAQ

How much does Warp's AI cost?

Warp's terminal is free, and its AI is metered in credits. As of October 1, 2026 Warp's pricing page lists Free at $0 with no bundled agent usage, Build at $20 a month with 1,500 credits, Max at $200 a month with 18,000 credits, Business at $50 per seat per month with 1,500 credits per seat for up to 25 seats, and Enterprise on custom terms; paying annually takes 10% off. Extra credits come in Reload packs from 400 credits for $10 to 6,500 for $100 on paid plans, and Warp says Free-plan Reload rates are 20% higher. Using your own API key or a custom inference endpoint for a model means that model's calls consume no Warp AI credits.

What is a Warp credit, and how many tokens is it?

A credit is Warp's unit of agent work, and Warp does not publish a conversion to tokens. Its documentation says credit usage scales with tokens processed and generally tracks the model's API pricing, but that there is no exact formula and two similar prompts can use different numbers of credits. It also meters three separate buckets from one balance: AI credits for the model call, compute credits for Warp-hosted sandboxes on cloud runs, and platform credits billed by the agent hour. To learn your own ratio, run a task and read the credit chip under the response in the app, or /cost in the Warp Agent CLI.

Does bringing my own API key or endpoint make Warp's AI free?

It removes AI credits for the requests you route through it, and not every other charge. Warp's Auto models always use Warp credits, custom routers cannot resolve to an endpoint-routed model, cloud agent runs always consume credits, and features such as Codebase Context rely on Warp's infrastructure. On Business and Enterprise, local agent runs that use your own key or endpoint still consume platform credits, billed by the agent hour. Warp's terms also limit self-serve BYOK and custom endpoints to individuals and organizations of 10 or fewer employees; larger organizations need Business or Enterprise.

Do unused Warp credits roll over?

Plan credits do not; Reload credits do. Warp's pricing FAQ says included monthly credits reset every 30 days from your subscription or renewal date, and its add-on credit documentation says purchased credits roll over across billing cycles and stay valid for 12 months from purchase, including after a move to the Free plan. On a team, purchased credits go into one shared pool that every member draws from once their own plan credits are used, so one heavy user can spend credits another member bought; admins can set a team-wide monthly spend cap.

Can Warp use an OpenAI-compatible endpoint like Kunavo?

Warp documents a custom inference endpoint for any service that implements the OpenAI Chat Completions API, configured in the Warp app under Settings by searching for inference endpoint, with the base URL, an API key and the model identifiers. The endpoint must be reachable at a public URL because the request is assembled on Warp's servers, and Warp states it cannot enforce Zero Data Retention for that traffic. Custom endpoints are available only in the Warp app, not the Warp Agent CLI. Kunavo's endpoint is public and OpenAI-compatible, but nobody at Kunavo has run Warp against it, so treat the first request as your own test.

Read on October 1, 2026: warp.dev/pricing (plan cards and the FAQ in its structured data) and Warp's documentation for Plans and billing, Credits, Add-on credits, Platform credits, Custom inference endpoint and Bring Your Own API Key. Not published anywhere read: a platform-credit rate per agent hour, and any conversion from credits to tokens. Not tested: Warp was not installed or subscribed, and no request went from Warp to Kunavo. Kunavo token rates come from the live catalog; every dollar figure in the token table is illustrative arithmetic.