Back to guides
Prices·June 18, 2026·Updated September 12, 2026·8 min read

Gemini API Pricing 2026 — Model Costs, Examples, and Cheaper OpenAI-Compatible Access

Gemini is among the best value-for-money cutting-edge AI models. See the current rates by model — about 70% below Google’s list price — with cost examples and the cheapest way to call Gemini in production.

Gemini offers some of the best value in frontier AI, and Kunavo provides it around 70% below Google's list price behind a single OpenAI-compatible API. This guide gives current rates by model, cost examples you can verify yourself, and the cheapest way to call Gemini in production.

Gemini API pricing at a glance

Rates are per 1M tokens in USD, as billed on Kunavo. The “Google list” column shows Google's published rate for the same model so you can see the difference.

ModelInput / 1MOutput / 1MGoogle list (input / output)Your savings
gemini-2-5-flash$0.09$0.75$0.30 / $2.50~70%
gemini-3-1-pro$0.70$4.20$2.00 / $12.00~65%

Flash is the workhorse for high-volume workloads; Pro is for harder reasoning, vision, and long-context tasks. Live rates are always shown on the pricing page and each model's page (gemini-2-5-flash, gemini-3-1-pro).

How Gemini token pricing works

You pay for input tokens (everything you send — system prompt, retrieved context, the user's message) and output tokens (what the model generates). Output is the more expensive side, so the biggest lever on a Gemini bill is how much text you let the model write. Images and audio are converted into token equivalents and billed on the same meter.

Worked cost examples

Real numbers at Kunavo's Gemini 2.5 Flash rate, except the last row, which uses Gemini 3.1 Pro:

WorkloadTokens (input / output)ModelCost
Chatbot turn1.000 / 300Flash$0.0003
RAG answer8.000 / 500Flash$0.0011
Batch classification (per document)500 / 20Flash$0.00006
Long-context analysis20.000 / 2.000Pro$0.0224

At these rates, a batch of 100,000 documents classified with Flash costs about $6, and one million chatbot turns cost about $315. Here's the calculation, ready to run:

gemini_custo.py
# Tarifas do Kunavo para o Gemini 2.5 Flash (USD por 1M de tokens)
IN_RATE, OUT_RATE = 0.09, 0.75

def custo(tokens_entrada: int, tokens_saida: int) -> float:
    return tokens_entrada / 1_000_000 * IN_RATE + tokens_saida / 1_000_000 * OUT_RATE

print(custo(1_000, 300))            # uma rodada de chatbot   -> $0.000315
print(custo(8_000, 500))            # uma resposta de RAG     -> $0.001095
print(custo(500, 20) * 100_000)     # lote de 100k documentos -> ~$6.00

Kunavo pricing and Stripe billing

There is no subscription or Google Cloud project. You add balance to a wallet (via Stripe or local payment methods), and calls are deducted at the token rates above. Pay-as-you-go starts with a minimum $10 deposit, the balance never expires, and larger deposits earn bonus credit. One wallet covers Gemini and all other models — Claude, GPT, image, video, and audio — so you don't have to reconcile a separate bill for each provider.

Which Gemini model should I choose?

  • gemini-2-5-flash — the default for chat, extraction, classification, summarization, and most RAG. Fast and the cheapest capable option.
  • gemini-3-1-pro — use it when Flash isn't precise enough: multi-step reasoning, code, vision, and very long context.

A good default is to route by difficulty: Flash for the common case, moving up to Pro only when a check fails. See the AI cost optimization guide for the routing pattern in code.

How to cut your Gemini bill

  1. Use a lower tier. Send the easy 80% to Flash; reserve Pro for the hard 20%.
  2. Limit output. Set max_tokens and stop sequences — output is the expensive side of the meter.
  3. Trim input. Retrieve fewer, better RAG chunks instead of cramming your whole knowledge base into the context.
  4. Batch requests. Group independent calls to keep latency low and avoid retry storms.

Frequently asked questions

Is the Gemini API free?

Google AI Studio offers a rate-limited free tier for prototyping. In production, you pay per token. Kunavo is pay-as-you-go starting with a minimum $10 top-up — you pay per token at the rates below, your balance never expires, and no Google Cloud billing account is required.

How much does Gemini 2.5 Flash cost?

On Kunavo, Gemini 2.5 Flash costs $0.09 per 1M input tokens and $0.75 per 1M output tokens — around 70% below Google's list price of $0.30 / $2.50. A typical chatbot turn (1K input, 300 output) costs about $0.0003.

Is Gemini cheaper than Claude or GPT?

Gemini 2.5 Flash is one of the cheapest capable models available — below Claude Haiku and most GPT tiers for high-volume workloads. Compare the full table on the pricing page.

How can I reduce Gemini API costs?

Use Flash, limit output, trim retrieved context, and batch requests. See the cost optimization guide for details. To start calling Gemini, see how to get a Gemini API key.