Gemini offers one of the best value propositions in frontier AI, and Kunavo serves it about 70% below Google's list price through a single OpenAI-compatible API. This Gemini API pricing guide gives current rates by model, cost examples you can recalculate yourself, and the cheapest way to call Gemini in production.
Gemini pricing at a glance
Rates are in USD per million tokens, as billed by Kunavo. The “Google price” column shows Google's published rate for the same model, for a direct comparison.
| Model | Input / 1M | Output / 1M | Google price (input / output) | Savings |
|---|---|---|---|---|
gemini-2-5-flash | $0.09 | $0.75 | $0.30 / $2.50 | ~70 % |
gemini-3-1-pro | $0.70 | $4.20 | $2.00 / $12.00 | ~65 % |
Flash is the workhorse for high volumes; Pro is for more challenging reasoning, vision, and long context. Live rates are always on the pricing page and on each model page (gemini-2-5-flash, gemini-3-1-pro).
How Gemini token pricing works
You pay for input tokens (everything you send—system prompt, retrieved context, user message) and output tokens (what the model generates). Output is the more expensive side, so the biggest lever on a Gemini bill is how much text you let the model write. Images and audio are converted into token equivalents and billed on the same meter.
Detailed cost examples
Real figures at Kunavo's Gemini 2.5 Flash rate, except for the last row, which uses Gemini 3.1 Pro:
| Workload | Tokens (input / output) | Model | Cost |
|---|---|---|---|
| Chatbot turn | 1,000 / 300 | Flash | $0.0003 |
| RAG response | 8,000 / 500 | Flash | $0.0011 |
| Batch classification (per document) | 500 / 20 | Flash | $0.00006 |
| Long-context analysis | 20,000 / 2,000 | Pro | $0.0224 |
At these rates, a batch of 100,000 documents classified with Flash costs about $6, and one million chatbot turns cost about $315. The calculation is executable as written:
# Tarifs Kunavo pour Gemini 2.5 Flash (USD par 1M de tokens)
IN_RATE, OUT_RATE = 0.09, 0.75
def cout(in_tokens: int, out_tokens: int) -> float:
return in_tokens / 1_000_000 * IN_RATE + out_tokens / 1_000_000 * OUT_RATE
print(cout(1_000, 300)) # un tour de chatbot -> $0.000315
print(cout(8_000, 500)) # une réponse RAG -> $0.001095
print(cout(500, 20) * 100_000) # lot de 100k documents -> ~$6.00Kunavo pricing and Stripe billing
No subscription, no Google Cloud project. You top up a balance (card, Apple Pay, Google Pay, Link), and calls are deducted at the per-token rates above. Pay-as-you-go starts with a minimum top-up of $10; the balance never expires, and larger top-ups provide bonus credit. Prices are in USD and converted to EUR by Stripe at checkout (the rate includes a 2–4% conversion fee paid by you; paying in USD avoids it); for B2B, reverse charge applies with an intra-community VAT number. One balance covers Gemini and all other models — Claude, GPT, image, video, and audio — with no provider-by-provider invoices to reconcile.
Which Gemini model should you choose?
- gemini-2-5-flash — the default choice for chat, extraction, classification, summarization, and most RAG. Fast and the cheapest capable option.
- gemini-3-1-pro — when Flash isn't accurate enough: multi-step reasoning, code, vision, and very long context.
A good pattern is difficulty-based routing: use Flash by default and escalate to Pro only when a check fails. The implementation is in the AI cost optimization guide.
How to reduce your Gemini bill
- Use a lower tier. Send the 80% of tasks that are simple to Flash; reserve Pro for the 20% that are difficult.
- Cap output. Set
max_tokensand stop sequences—output is the more expensive side of the meter. - Trim input. Retrieve fewer, more relevant RAG chunks instead of cramming the entire knowledge base into the context.
- Batch requests. Batch independent calls to keep latency low and avoid retry storms.
The same key also calls Google's video models—the per-clip price is detailed in Veo 3 pricing.
Frequently asked questions
Is the Gemini API free?
Google AI Studio offers a rate-limited free tier for prototyping. In production, you pay per token. Kunavo is pay-as-you-go starting with a minimum $10 top-up — you pay per token at the rates below, your balance never expires, and no Google Cloud billing account is required.
How much does Gemini 2.5 Flash cost?
On Kunavo, Gemini 2.5 Flash costs $0.09 per 1M input tokens and $0.75 per 1M output tokens — around 70% below Google's list price ($0.30 / $2.50). A typical chatbot turn (1K input, 300 output) costs about $0.0003.
Is Gemini cheaper than Claude or GPT?
Gemini 2.5 Flash is one of the cheapest capable models available — below Claude Haiku and most GPT tiers for high-volume workloads. Compare the full table on the pricing page, and Claude rates in the Claude API pricing guide.
How can I reduce Gemini API costs?
Use a lower tier, limit output, trim retrieved context, and batch requests. See the cost optimization guide for details. To start calling Gemini, see how to get a Gemini API key.