Back to guides
Pricing·October 1, 2026·Updated October 3, 2026·9 min read

Qwen Code pricing: Coding Plan, Token Plan, and pay-as-you-go API

The client is free. The free OAuth offering no longer exists. Everything else is purchased inference, in four different units.

Qwen Code itself is free: the client is open source under Apache-2.0, with no paid tier, so every euro of a Qwen Code budget goes toward model inference, purchased separately. The free Qwen OAuth offer still described by most tutorials was discontinued on April 15, 2026. In 2026, a budget therefore starts with a paid endpoint: an Alibaba Cloud Model Studio Coding Plan, a Credits Token Plan package, pay-as-you-go token billing, or a third-party endpoint you configure yourself.

These four options are measured in four different units — requests, Credits, tokens, and whatever your own provider charges — which is why comparing the entry price gives the wrong answer. This page separates software cost from inference cost, provides a recalculable estimate, and flags what could not be verified. The detailed English version is Qwen Code pricing.

The software is free; the models are not

The QwenLM/qwen-code repository is under Apache-2.0 and actively developed; the GitHub API listed it as not archived on October 1, 2026, with a push that same day. Releases come out at least weekly: v0.24.7 was published on npm (@qwen-code/qwen-code) on September 29, 2026. Installation is available via standalone script, npm (Node.js 22 or newer), or Homebrew. The same free repository provides a desktop app, web interface (qwen serve --open), VS Code, Zed, and JetBrains extensions, and TypeScript, Python, and Java SDKs; none of these surfaces is a paid option, and none includes inference.

One connection to set aside: Qwen Code was initially based on Google Gemini CLI v0.8.2, but the project has not synchronized with it since v0.1. Gemini CLI quotas and prices do not apply.

The free limit: what ended, and when

The authentication documentation states that the free Qwen OAuth offer was discontinued on 2026-04-15 and that already-cached tokens may work briefly, but new requests are rejected. In the interface, selecting Qwen OAuth now displays “Discontinued — switch to Coding Plan or API Key”.

PeriodFree OAuth quotaSource
Through v0.9.0 (npm, February 3, 2026)2,000 requests/day, 60 requests/minuteAuthentication documentation at the v0.9.0 tag
Starting with v0.10.0 (npm, February 9, 2026)1,000 requests/day, 60 requests/minuteSame file at tags v0.10.0 through v0.14.0
Starting April 15, 2026None — discontinuedCurrent documentation; issue #3316

The exact date when the limit changed from 2,000 to 1,000 requests was not announced anywhere: the range above comes from the version tag where the documentation changed. The figure of 100 requests per day that circulates does not appear in any of the verified tags; consider it unsupported. Qwen OAuth models are also hard-coded and cannot be redirected via modelProviders.

Three official endpoints, three billing units

In /auth, Alibaba ModelStudio opens a submenu for Coding Plan, Token Plan, and Standard API Key. These are not three ways to pay the same bill: the documentation gives each its own host and key, and requires the baseUrl to match the key’s plan. Coding Plan keys begin with sk-sp-.

RouteBilled inPublished price (October 1, 2026)Host
Coding Plan Pro (international)Requests$50/monthcoding-intl.dashscope.aliyuncs.com
Coding Plan Pro (Chinese site)Requests¥200/month; ¥39.90 for the first month for a new customercoding.dashscope.aliyuncs.com
Token Plan, PersonalCreditsLite $8 (promo $6), Essential $16 ($10), Standard $25 ($18), Pro $80 ($68) per monthtoken-plan.ap-southeast-1.maas.aliyuncs.com
Token Plan, TeamCreditsStandard $30 ($20), Pro $100 ($75), Max $200 per seat per monthSame Singapore host
Standard API KeyTokensPer model, by input-length tierdashscope.aliyuncs.com

The Coding Plan page applies three simultaneous limits to Pro: up to 6,000 requests per 5 hours, 45,000 per week, and 90,000 per month; once any one is reached, calls are suspended. The 5-hour quota is released on a rolling basis, the weekly quota resets Monday at 00:00 (UTC+8), and the monthly quota resets on the renewal date. On October 1, 2026, the page also reported limited availability: a limited number of places, first come first served, replenished daily (00:00 UTC+8 internationally, 09:30 UTC+8 on the Chinese site). The Lite offer no longer exists: new subscriptions stopped on March 20, 2026, and renewals and upgrades stopped on April 13, 2026. And ¥200 is not the conversion of $50: treat the two storefronts as separate offers.

What a request really costs: each question consumes quota based on the number of model calls, and Alibaba indicates 5 to 10 calls for a simple task, 10 to 30 or more for a complex task. One instruction therefore amounts to several requests; at that rate, 90,000 requests per month represent roughly 3,000 to 18,000 tasks—but Alibaba presents these figures as typical, not guaranteed. Measure on your account before sizing.

The Token Plan calls for two remarks. First, it has changed: on October 1, 2026, the Token Plan page assigned Personal plans a monthly quota (30-day cycle starting from subscription) of 11,500, 25,500, 45,000, and 180,000 Credits for Lite, Essential, Standard, and Pro, suspended the service once the quota was reached, sold additional packages at $15 for 20,000 Credits, and specified that unused quota does not roll over. Second, it publishes no Credits → tokens conversion or model multiplier: nothing makes it possible to say how many tokens $18 per month buy. The offer is sold only in the Singapore region.

Pay-as-you-go, per token

The Model Studio pricing page gives prices in dollars per million tokens, by input-length tier for each request. Singapore column below; China (Beijing) rates are separate and significantly lower, so check your billing region.

ModelInput 0-32K32K-128K128K-256K256K-1M
qwen3-coder-plus$1 input / $5 output$1.8 / $9$3 / $15$6 / $60
qwen3-coder-next$0.3 input / $1.5 output$0.5 / $2.5$0.8 / $4Not listed

Read the final tier before budgeting a long agent session: qwen3-coder-plus output rises from $5 to $60 per million as soon as a request's input exceeds 256K. New Singapore accounts receive 1 million free tokens per model, valid for 90 days—a trial, not a free offer. Also beware of the official qwen3-coder-plus price reduction page: the cache discount it announces ran from July 24 to August 23, 2025.

Connect Qwen Code to another endpoint

Qwen Code supports four protocols—OpenAI Chat Completions, OpenAI Responses, Anthropic, and Gemini/Vertex—and each accepts a custom base URL; its own documentation cites OpenRouter and Fireworks AI as alternatives after OAuth was discontinued. Kunavo serves /v1/chat/completions, /v1/responses, and /v1/messages under https://api.kunavo.com/v1.

Let's be clear about what this provides. The Kunavo catalog contains no Qwen text models: this is not a cheaper way to run Qwen, but a way to run Claude or GPT in Qwen Code with a single prepaid balance. Kunavo has also not run Qwen Code against these endpoints; the configuration below comes from the Qwen Code documentation, as does the English Qwen Code configuration page.

To be merged into ~/.qwen/settings.json
{
  "modelProviders": {
    "openai": [
      {
        "id": "claude-sonnet-5",
        "name": "Claude Sonnet 5",
        "baseUrl": "https://api.kunavo.com/v1",
        "envKey": "KUNAVO_API_KEY"
      }
    ]
  },
  "security": {
    "auth": {
      "selectedType": "openai"
    }
  },
  "model": {
    "name": "claude-sonnet-5"
  }
}

Four facts that save hours of debugging, taken from the model providers reference: the baseUrl stops at /v1 (the SDK adds the path; with /v1/chat/completions, it returns a 404); the key is read from the variable named by envKey, preferably via .qwen/.env rather than in plain text in settings.json; the selected modelProviders entry takes precedence over the CLI options --openai-base-url and --openai-api-key, which is why they may appear to be ignored; and wireApi: "responses" has neither endpoint detection nor automatic fallback—start with Chat Completions.

Two limitations to plan for. The built-in web_search tool relies on DashScope's server-side search: active with a Standard API Key or a Token Plan, inactive with the Coding Plan (not verified on this endpoint), and inactive with a third-party provider or an endpoint on another host. The documented alternative is an MCP search server. And a Kunavo route serves chat only: no embeddings, summarization, or speech recognition.

A recalculable monthly estimate

Suppose a month of 20 million uncached input tokens and 1 million output tokens. At the current Kunavo catalog rates:

ModelInput / output per millionEstimated monthPossible role
Claude Haiku 4.5$0.70 / $3.50$17.50Bounded changes and high volume
GPT-5.6 Terra$0.70 / $4.20$18.20Mixed reasoning and tools
Claude Sonnet 5$1.40 / $7.00$35.00Difficult repository changes

This is arithmetic based on an assumed workload, not a measured Qwen Code task or a billing cap; cache, external tools, and taxes are excluded, and it assumes the cheaper model completes the work without another attempt. The Kunavo catalog amount is a billing floor, not a ceiling: when the upstream reports its cost, the bill is the higher of the catalog cost and the upstream cost multiplied by the applicable markup. This table is not comparable to the $50 Coding Plan (which counts calls, not tokens) or the Token Plan (Credits without a published conversion): to compare a plan with a meter, you need your own measured usage on both sides.

Kunavo starts with a prepaid top-up of $10, without a subscription — the amount required to open a funded account, not the price of a task. Stripe checkout accepts cards (Visa, Mastercard, American Express), Apple Pay, Google Pay, and Link; SEPA Direct Debit is currently not offered.

Which route to choose, and when

ChooseWhen it fitsWhat you give up
Coding PlanDaily, regular Qwen usage that remains below the request limitsCalls suspended at each limit; no integrated web search; quota shared with any other tool using the key; limited seats
Pay-as-you-go per tokenIrregular or occasional usage, with long sessions to price preciselyInput-length tiers that drive up the rate for large requests
Token Plan CreditsYou have confirmed the Credits rate for your models in your consoleNo public conversion, monthly quota that suspends the service once exhausted, Singapore only
A gateway such as KunavoYou want Claude or GPT in Qwen Code with a single prepaid balanceNo Qwen models; manual configuration; no integrated web search
Local inferenceYou have suitable hardware and the model handles your toolsHardware, maintenance, and tool-call reliability become your responsibility

Keep the two questions separate: the price per million tokens is found on one page; the lowest cost to complete your task is a measurement that nobody has published for this client, including us. Run the same three tasks—a small modification, a bug fix, and a multi-file investigation—through two routes and compare the recorded amounts. To try the gateway route, create a Kunavo account and then follow the quickstart.

Frequently asked questions

How much does Qwen Code cost?

The Qwen Code client costs nothing: it is open source under Apache-2.0 on GitHub (QwenLM/qwen-code), with no paid tier, seat fee, or subscription, and the desktop app, web interface, editor extensions, and SDKs are in the same free repository. What you pay for is model inference, purchased separately from the endpoint you configure.

Is Qwen Code still free? What is the free limit?

The software remains free, but free hosted inference does not. Qwen Code’s authentication documentation states that the free Qwen OAuth offer was discontinued on April 15, 2026, and OAuth is no longer selectable in the /auth menu. Tutorials claiming 2,000 requests per day describe the quota provided through v0.9.0 in February 2026. On Alibaba’s side, new accounts still receive trial credit: 1 million tokens per model in the Singapore region, valid for 90 days — a trial, not a free offer.

How much does Alibaba’s Coding Plan for Qwen Code cost?

On the international site, Coding Plan Pro costs $50 per month, with three simultaneous limits: 6,000 requests per 5-hour period, 45,000 per week, and 90,000 per month; once a limit is reached, calls are suspended. A request corresponds to one model call, and Alibaba says a simple task generally consumes 5 to 10, while a complex task consumes 10 to 30 or more. On the Chinese site, the same Pro plan costs ¥200 per month (¥39.90 for the first month for a new customer). On October 1, 2026, the offer had limited availability, with places replenished daily. The former Lite offer is no longer sold.

What is the cheapest API for Qwen Code?

The lowest price per million tokens and the lowest cost to complete a task are two different questions, and only the first can be read from a pricing table. Among Alibaba’s published prices for Singapore, qwen3-coder-next at $0.30 input and $1.50 output per million (0–32K tier) is cheaper than qwen3-coder-plus at $1 and $5. Whether a cheaper model is actually cheaper on your repository depends on how many attempts it needs, which no pricing table measures.

Can I pay for Qwen models with Kunavo credit?

No. The Kunavo catalog contains no Qwen text model, so Kunavo is not a cheaper route to Qwen. What Kunavo can do is serve Claude and GPT models through OpenAI-compatible endpoints (Chat Completions and Responses) and Anthropic Messages, formats that Qwen Code can use. Kunavo has not run Qwen Code against these endpoints: this is a documented configuration to try, not a verified integration.

Verified October 1, 2026: Qwen Code repository, versions, and documentation (main branch and v0.9.0 to v0.15.0 tags for free-quota history), npm registry, Coding Plan pages (international and Chinese site), Token Plan, and Model Studio pricing. Kunavo rates come from the live catalog; monthly totals are illustrative token arithmetic, not measured task costs.