Back to guides
Prices·September 11, 2026·Updated October 3, 2026·10 min read

OpenAI API costs 2026—what GPT really costs per request

The OpenAI API bills by token; a ChatGPT subscription is separate. Current per-model rates, what a request really costs—and the two pitfalls that derail every estimate.

The OpenAI API is billed per token, not per month. Input and output tokens have separate rates, and output is the more expensive part. On Kunavo, rates range from $0,04 / $0,20 per 1M tokens for GPT-6 Luna to $4,00 / $20,00 for GPT-6 Astra, OpenAI's latest flagship; in between are, for example, GPT-6 Sol at $0,80 / $4,00 and the workhorse GPT-5.6 Terra at $0,70 / $4,20 — depending on the model, 60–65% below OpenAI's list price. A ChatGPT subscription is separate: it includes no API usage, and the API requires no subscription.

Rates as of October 3, 2026. This page reads Kunavo rates live from the model catalog, so the figures shown here are what the API bills. The “OpenAI list” column shows OpenAI's published price for the same model; the source is openai.com/api/pricing.

ChatGPT subscription and OpenAI API: two products, two bills

People searching for “ChatGPT API cost” often first want to know something else: whether their existing ChatGPT subscription already covers the API. The answer is no.

What you useHow billing worksFor whom
ChatGPT (Free, Plus, Pro …)Fixed monthly fee, with the plan's usage limitsPeople who work in the chat window
OpenAI API (key + balance)Per token, no monthly feeYour own software, automation, and agents

A subscription includes neither an API key nor API balance and does not reduce the API bill by even a cent; conversely, API balance does not increase ChatGPT limits. OpenAI lists subscription prices at openai.com/chatgpt/pricing; they change often, so we do not include them here. This page covers only the API — the part billed per token.

OpenAI API prices by model

Rates in USD per million tokens, as billed by Kunavo. The “OpenAI list” column shows the published price for the same model.

ModelInput / 1MOutput / 1MOpenAI list price (input / output)Savings
gpt-6-astra$4,00$20,00$10,00 / $50,00~60%
gpt-5-6-sol$2,00$12,00$5,00 / $30,00
Promotional price $4,00 / $20,00 (at least until November 21, 2026)
~60%
~40% below the promotional price
gpt-5-5$2,00$12,00$5,00 / $30,00~60%
gpt-5-6-terra$0,70$4,20$2,00 / $12,00~65%
gpt-6-sol$0,80$4,00$2,00 / $10,00~60%
gpt-6-luna$0,04$0,20$0,10 / $0,50~60%

Live rates are always listed on the Pricing page and on each model page, such as GPT-6 Astra, GPT-6 Sol, or GPT-6 Luna. Which model to use for what:

  • gpt-6-astra — OpenAI's latest flagship (since September 3, 2026): top-tier reasoning and long agentic coding runs. For the difficult minority of tasks.
  • gpt-6-sol — since September 22, 2026; OpenAI positions it for complex coding and agentic workflows. On Kunavo, $0,80 / $4,00, with output priced below GPT-5.6 Terra, slightly more for input. The sensible default for coding and agents.
  • gpt-6-luna — since September 22, 2026; according to OpenAI, its most efficient model for narrow, high-volume tasks such as classification and extraction. At $0,04 / $0,20, it is the lowest-cost tier on Kunavo.
  • gpt-5-6-sol — top-tier reasoning and agentic coding from the 5.6 generation; for prompts already tuned for it.
  • gpt-5-6-terra — the workhorse of the 5.6 family: RAG, assistants, classification, extraction, and most production tasks.
  • gpt-5-5 — the flagship of the previous generation. GPT-5.6 Sol costs the same rate on Kunavo ($2,00 / $12,00); anyone still running gpt-5-5 can switch to the newer generation by changing the model name, with no additional cost per token.

GPT-6 Sol and GPT-6 Luna. On September 22, 2026, OpenAI expanded the GPT-6 family with two models that it says build on the advances behind GPT-6 Astra. OpenAI lists GPT-6 Sol at $2,00 / $10,00 and GPT-6 Luna at $0,10 / $0,50 per 1M tokens; on Kunavo, they cost $0,80 / $4,00 and $0,04 / $0,20. Both have a 1.050.000-token context window and support up to 128.000 output tokens.

What GPT-6 Astra costs extra. Compared with GPT-6 Sol, GPT-6 Astra costs 5 times as much for both input and output; compared with GPT-5.6 Sol, it costs 2 times as much for input and 1,7 times as much for output. The honest recommendation: use GPT-6 Sol as the default for coding and agents, GPT-6 Luna for high-volume tasks, and Astra selectively where a comparison using the same inputs shows that Sol is not enough — the model name is a single word in the request, so testing is cheap.

A surcharge for very long requests. Above 272,000 prompt tokens, GPT-5.5, the GPT-5.6 family, and all three GPT-6 models — Astra, Sol, and Luna — bill the entire request at 2× the input rate and 1.5× the output rate. If you include entire codebases or long documents in one call, be aware of this threshold: it can have a greater effect than model choice.

What a request really costs

Calculated at Kunavo rates, without caching or reasoning tokens — so these are the lower bounds of what will appear on the bill:

TaskTokens (input / output)GPT-6 LunaGPT-5.6 TerraGPT-6 SolGPT-5.6 SolGPT-6 Astra
Short chat session1.000 / 300$0,0001$0,00196$0,002$0,0056$0,01
RAG-Antwort6.000 / 500$0,00034$0,0063$0,0068$0,018$0,034
Agentic coding step25.000 / 1.200$0,00124$0,0225$0,0248$0,0644$0,124
Heavy reasoning task20.000 / 3.000$0,0014$0,0266$0,028$0,076$0,14
Batch: 100,000 classifications500 / 20 per request$2,40$43,40$48,00$124,00$240,00

With the minimum $10 top-up, you can run about 1.590 RAG responses on GPT-5.6 Terra, or about 29.400 on GPT-6 Luna. The output rate for GPT-6 Luna is about 100 times that of GPT-6 Astra — for simple tasks, model choice is the biggest single cost lever. To calculate using your own numbers:

openai_kosten.py
# OpenAI-Tarife auf Kunavo (USD pro 1M Tokens): (Input, Output)
RATES = {
    "gpt-6-astra":   (4, 20),
    "gpt-5-6-sol":   (2, 12),
    "gpt-5-5":       (2, 12),
    "gpt-5-6-terra": (0.7, 4.2),
    "gpt-6-sol":     (0.8, 4),
    "gpt-6-luna":    (0.04, 0.2),
}

def kosten(model: str, frisch: int, output: int, cache_treffer: int = 0) -> float:
    """frisch = nicht gecachter Input; output schließt Reasoning-Tokens ein."""
    i, o = RATES[model]
    return (
        frisch / 1_000_000 * i
        + output / 1_000_000 * o
        + cache_treffer / 1_000_000 * i * 0.10  # Cache-Treffer: 10 % des Input-Tarifs
    )

print(kosten("gpt-5-6-terra", 1_000, 300))    # kurze Chat-Runde
print(kosten("gpt-5-6-terra", 6_000, 500))    # RAG-Antwort
print(kosten("gpt-5-6-sol", 25_000, 1_200))   # agentischer Coding-Schritt
print(kosten("gpt-6-sol", 25_000, 1_200))     # derselbe Schritt auf GPT-6 Sol
print(kosten("gpt-6-astra", 20_000, 3_000))   # schwere Reasoning-Aufgabe
print(kosten("gpt-6-luna", 500, 20))          # eine Klassifizierung

# Ein Monat auf Sol: 3 Mio. frische und 12 Mio. gecachte Input-Tokens,
# 1 Mio. Output-Tokens (Cache-Writes nicht eingerechnet)
print(kosten("gpt-5-6-sol", 3_000_000, 1_000_000, cache_treffer=12_000_000))

Two cost traps every estimate misses

First: reasoning tokens. Reasoning models like GPT-6 Astra or GPT-5.6 Sol think through the problem using internal tokens before answering. These do not appear in the response text, but are billed at the output rate. The estimates on this page do not include them, so the actual bill for difficult tasks will be higher. The actual count appears in every response under usage (completion_tokens_details.reasoning_tokens).

Second: retries. A response that is returned successfully but discarded by your code — due to the wrong format, for example, or because it is retried — is still paid for. With agents, one visible task often means a dozen billable runs, which is where the gap between an estimate and the bill is greatest. Requests that fail on Kunavo's side are not charged.

Prompt caching

If your application sends the same long system prompt or documents with every request, cache hits are billed at 10% of the input rate — on GPT-5.6 Sol, for example, that is $0,20 instead of $2,00 per 1M tokens. On Kunavo, writing to the cache costs 1.25× the input rate for GPT-5.6 Sol, Terra, and GPT-6 Astra, and the standard input rate for GPT-5.5. See the caching documentation for how to get a cache hit.

Switching existing OpenAI code

Kunavo is OpenAI-compatible: change the base_url and key, and nothing else. chat.completions and responses keep their format, and the response includes the familiar usage field, which lets you trace every charge.

openai_kunavo.py
from openai import OpenAI

client = OpenAI(
    api_key="sk-kn-...",
    base_url="https://api.kunavo.com/v1",  # statt api.openai.com
)

resp = client.chat.completions.create(
    model="gpt-5-6-terra",
    messages=[{"role": "user", "content": "Fasse diese Rechnung in drei Sätzen zusammen."}],
)
print(resp.choices[0].message.content)
print(resp.usage)  # Input-, Output- und Reasoning-Tokens: die Grundlage der Abrechnung

Create your key (sk-kn-…) in the dashboard after signing up. There, you can also set a monthly spending limit for each key — once reached, the key returns 402 — and an IP allowlist. The steps are in the Quickstart.

Paying for the OpenAI API from Germany

Kunavo bills through Stripe using prepaid balance; no subscription or OpenAI Platform account setup is required. The checkout offers:

  • Cards: Visa, Mastercard, American Express, JCB, and UnionPay — Visa and Mastercard debit cards work like credit cards
  • Apple Pay, Google Pay, and Link

SEPA Direct Debit is currently not offered. Top-ups are denominated in US dollars; Stripe displays the amount in your currency at checkout. The minimum top-up is $10; larger top-ups include bonus credit: $100 becomes $110.00, $1,000 becomes $1,200.00, and $5,000 becomes $6,250.00. The balance never expires, and failed requests are not charged. Details are available in the Billing documentation.

Honestly: when buying directly from OpenAI is a better fit

Kunavo is a gateway with shared capacity: there is no quota dedicated to your account and no contractual SLA. If you need a guaranteed quota or an SLA, for example for a product with a promised availability level, you should buy directly from OpenAI. This page does not claim to cover that case.

And for people who work in a chat window all day, a ChatGPT subscription is often the cheaper choice. The calculation is simple division: subscription price ÷ cost per chat round = break-even point. A round with 2,000 input and 700 output tokens costs about $0,0124 on GPT-5.6 Sol, or about $0,00434 on GPT-5.6 Terra. If you have more rounds than that calculation yields month after month, and stay within the plan limits, the subscription is the better deal. The API is better for software, automation, and variable usage — when you want a quiet month to cost nothing.

Lower OpenAI API costs—in this order

  1. Choose the model for the task. Leaving the largest model as the default is the most common cause of a surprising bill. Use gpt-6-luna for classification, extraction, and other high-volume tasks; gpt-6-sol for coding and agents; and Astra for the difficult minority.
  2. Cap output. Output costs several times more than input, and reasoning counts toward it—so a limit over max_tokens has a double effect.
  3. Use prompt caching. A long, stable system prompt costs one tenth of the input rate when it is a cache hit.
  4. Send less context. Use a few precise RAG passages instead of half the knowledge base—and stay below the 272,000-token threshold where possible.

Recalculate with your own numbers: the OpenAI API pricing calculator includes cached input and reasoning tokens, while the token calculator compares Claude and GPT side by side. For Anthropic's equivalent, see Claude pricing. If limits, rather than costs, are the bottleneck, see OpenAI API rate limits (in English).

Frequently asked questions

How much does the OpenAI API cost?

The OpenAI API is billed per token, with no monthly fee: input and output tokens have separate rates, and output is the more expensive part. On Kunavo, GPT-5.6 costs Terra $0,70 / $4,20 per 1M tokens (input / output), GPT-5.6 costs Sol $2,00 / $12,00, and OpenAI's latest flagship, GPT-6, costs Astra $4,00 / $20,00. OpenAI's list prices for the same models are $2,00 / $12,00, $5,00 / $30,00, and $10,00 / $50,00; for GPT-5.6 Sol, OpenAI currently charges a promotional price of $4,00 / $20,00, according to the pricing page at least until November 21, 2026. As of October 3, 2026.

Is the API included in a ChatGPT subscription?

No. A ChatGPT subscription such as Plus or Pro pays for using ChatGPT itself — in the browser, app, or desktop app. It does not include an API key or API balance, and does not reduce your API bill by even a cent. To call GPT from your own software, you need an API key and pay per token, separately from the subscription. Conversely, API balance does not increase your ChatGPT limits.

How much does a ChatGPT API request cost?

What people informally call the “ChatGPT API” is the OpenAI API with GPT models, and each request is billed by token. At Kunavo rates, a RAG response with 6,000 input and 500 output tokens on GPT-5.6 Terra costs around $0,0063; an agentic coding step with 25,000 / 1,200 tokens on GPT-5.6 Sol costs around $0,0644; and a heavy task with 20,000 / 3,000 tokens on GPT-6 Astra costs around $0,14. Reasoning tokens are charged at the output rate.

Can I try the OpenAI API for free?

Kunavo does not provide starting credits, but it also has no monthly fee or minimum term. The smallest way to get started is the $10 minimum top-up — enough for about 1.590 RAG responses on GPT-5.6 Terra. The balance never expires, a month with no calls costs nothing, and failed requests are not charged.

How are reasoning tokens billed?

At the Output rate. Reasoning models such as GPT-6 Astra or GPT-5.6 Sol think in internal tokens before answering; those tokens do not appear in the response text, but they are included in billing. Each response reports the actual number in the usage field (completion_tokens_details.reasoning_tokens). For difficult tasks, this is the most common reason why the actual bill exceeds an estimate; a limit on max_tokens caps it.

How much does GPT-6 Astra cost?

GPT-6 Astra has been OpenAI's latest flagship since September 3, 2026; OpenAI lists it at $10,00 / $50,00 per 1M tokens. On Kunavo, it costs $4,00 / $20,00 — around 60% below the list price. Compared with GPT-5.6 Sol, that is 2 times the input price and 1,7 times the output price; Astra is therefore worthwhile for the difficult minority of tasks where Sol demonstrably falls short. Above 272,000 prompt tokens, Astra also bills the entire request at 2x the input rate and 1.5x the output rate.

How much do GPT-6 Sol and GPT-6 Luna cost?

OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026: Sol for complex coding and agentic workflows, and Luna, in OpenAI's words, as its most efficient model for narrowly defined tasks at high volume. OpenAI lists Sol at $2,00 / $10,00 and Luna at $0,10 / $0,50 per 1M tokens (input / output); on Kunavo, they cost $0,80 / $4,00 and $0,04 / $0,20, beide rund 60% darunter. GPT-6 Luna is therefore the cheapest GPT model on Kunavo. As with GPT-5.5, the GPT-5.6 family, and GPT-6 Astra, a request with more than 272,000 prompt tokens is billed in full at 2x the input rate and 1.5x the output rate.

Can I pay by SEPA direct debit from Germany?

No, SEPA Direct Debit is currently not offered. Stripe checkout offers cards (Visa, Mastercard, American Express, JCB, UnionPay), Apple Pay, Google Pay, and Link; a Visa or Mastercard debit card works like a credit card. It is a prepaid balance with a $10 minimum top-up, without a subscription, and it never expires.

Will my existing OpenAI code work with Kunavo?

Yes. Kunavo is compatible with OpenAI: set base_url to https://api.kunavo.com/v1 and replace the key with a Kunavo key (sk-kn-…); chat.completions and responses retain their format. The model names are gpt-6-astra, gpt-5-6-sol, gpt-5-5, gpt-5-6-terra, gpt-6-sol and gpt-6-luna. The same key also accesses Claude and image and video models.