Gemini offers some of the best value in frontier AI, and Kunavo provides it around 70% below Google's list price behind a single OpenAI-compatible API. This Gemini API pricing guide gives current rates by model, cost examples you can verify yourself, and the cheapest way to call Gemini in production.
Gemini pricing at a glance
Rates are per 1M tokens in USD, as billed on Kunavo. The “Google list” column shows Google's published rate for the same model so you can see the difference.
| Model | Input / 1M | Output / 1M | Google list (input / output) | Your savings |
|---|---|---|---|---|
gemini-2-5-flash | $0.09 | $0.75 | $0.30 / $2.50 | ~70% |
gemini-3-1-pro | $0.70 | $4.20 | $2.00 / $12.00 | ~65% |
Flash is the workhorse for high-volume workloads; Pro is for harder reasoning, vision, and long-context tasks. Live rates are always shown on the pricing page and each model's page (gemini-2-5-flash, gemini-3-1-pro).
How Gemini token pricing works
You pay for input tokens (everything you send — system prompt, retrieved context, the user's message) and output tokens (what the model generates). Output is the more expensive side, so the biggest lever on a Gemini bill is how much text you let the model write. Images and audio are converted into token equivalents and billed on the same meter.
Worked cost examples
Real numbers at Kunavo's Gemini 2.5 Flash rate, except the last row, which uses Gemini 3.1 Pro:
| Workload | Tokens (input / output) | Model | Cost |
|---|---|---|---|
| Chatbot turn | 1.000 / 300 | Flash | $0.0003 |
| RAG response | 8.000 / 500 | Flash | $0.0011 |
| Batch classification (per document) | 500 / 20 | Flash | $0.00006 |
| Long-context analysis | 20.000 / 2.000 | Pro | $0.0224 |
At these rates, a batch of 100,000 documents classified with Flash costs about $6, and one million chatbot turns cost about $315. Here's the calculation, ready to run:
# Kunavo-Tarife für Gemini 2.5 Flash (USD pro 1M Tokens)
IN_RATE, OUT_RATE = 0.09, 0.75
def kosten(in_tokens: int, out_tokens: int) -> float:
return in_tokens / 1_000_000 * IN_RATE + out_tokens / 1_000_000 * OUT_RATE
print(kosten(1_000, 300)) # ein Chatbot-Turn -> $0.000315
print(kosten(8_000, 500)) # eine RAG-Antwort -> $0.001095
print(kosten(500, 20) * 100_000) # Batch mit 100k Docs -> ~$6.00Kunavo pricing and Stripe billing
There is no subscription or Google Cloud project. You add balance to a wallet (via Stripe or local payment methods), and calls are deducted at the token rates above. Pay-as-you-go starts with a minimum $10 top-up, the balance never expires, and larger top-ups earn bonus credit. One wallet covers Gemini and every other model — Claude, GPT, image, video, and audio — so you don't have to reconcile a separate bill for each provider.
Which Gemini model should I choose?
- gemini-2-5-flash — the default for chat, extraction, classification, summarization, and most RAG. Fast and the cheapest capable option.
- gemini-3-1-pro — use it when Flash isn't precise enough: multi-step reasoning, code, vision, and very long context.
A good pattern is to route by difficulty: Flash for the common case, moving up to Pro only when a check fails. See the AI cost optimization guide for the routing pattern in code.
How to cut your Gemini bill
- Use a lower tier. Send the easy 80% to Flash; reserve Pro for the hard 20%.
- Limit output. Set
max_tokensand stop sequences — output is the expensive side of the meter. - Trim input. Retrieve fewer, better RAG chunks instead of cramming your whole knowledge base into the context.
- Batch requests. Group independent calls to keep latency low and avoid retry storms.
The same key also calls Google's video models — see Veo 3 costs for per-clip pricing and OpenAI API costs for a comparison with OpenAI.
Frequently asked questions
Is the Gemini API free?
Google AI Studio offers a rate-limited free tier for prototyping. In production, you pay per token. Kunavo is pay-as-you-go starting with a minimum $10 top-up — you pay per token at the rates below, your balance never expires, and no Google Cloud billing account is required.
How much does Gemini 2.5 Flash cost?
On Kunavo, Gemini 2.5 Flash costs $0.09 per 1M input tokens and $0.75 per 1M output tokens — around 70% below Google's list price of $0.30 / $2.50. A typical chatbot turn (1K input, 300 output) costs about $0.0003.
Is Gemini cheaper than Claude or GPT?
Gemini 2.5 Flash is one of the cheapest capable models available — below Claude Haiku and most GPT tiers for high-volume workloads. Compare the full table on the pricing page.
How can I reduce Gemini API costs?
Use Flash, limit output, trim retrieved context, and batch requests. See the cost optimization guide for details. To start calling Gemini, see how to get a Gemini API key.