Back to guides
Price·June 18, 2026·Updated September 12, 2026·8 min read

Gemini API pricing 2026 — a complete guide to Gemini API fees

Gemini offers some of the best value among frontier AI models, and Kunavo provides it 30–70% below Google's list price depending on the model. Current rates by model, directly verifiable cost examples, and the cheapest way to call it in production, all in one place.

Gemini is among the best-value options in frontier AI, and Kunavo prices it about 70% below Google’s list price behind a single OpenAI-compatible API. This guide covers current rates by model, cost calculation examples you can verify yourself, and the cheapest way to call Gemini in production.

Gemini pricing at a glance

Rates are in USD per 1M tokens, as billed on Kunavo. The “Google list price” column shows Google’s published rate for the same model, included so you can see the difference.

ModelInput / 1MOutput / 1MGoogle list price (input / output)Savings
gemini-2-5-flash$0.09$0.75$0.30 / $2.50~70%
gemini-3-1-pro$0.70$4.20$2.00 / $12.00~65%

Flash is the workhorse for high-volume workloads; Pro is for harder reasoning, vision, and long-context tasks. Live rates are always available on the pricing page and each model page (gemini-2-5-flash, gemini-3-1-pro).

How Gemini token pricing works

You pay for input tokens (everything you send — system prompt, retrieved context, user message) and output tokens (what the model generates). Output is more expensive, so the biggest lever on your Gemini bill is how much text you let the model write. Images and audio are converted to token equivalents and billed on the same meter.

Cost calculation examples

Actual figures calculated using Kunavo’s Gemini 2.5 Flash rate, except the last row, which uses Gemini 3.1 Pro:

TaskTokens (input / output)ModelCost
Chatbot turn1,000 / 300Flash$0.0003
RAG response8,000 / 500Flash$0.0011
Batch classification (per document)500 / 20Flash$0.00006
Long-context analysis20,000 / 2,000Pro$0.0224

At these rates, classifying a batch of 100,000 documents on Flash costs about $6, and one million chatbot turns cost about $315. Here’s the runnable calculation:

gemini_cost.py
# Kunavo Gemini 2.5 Flash 요율 (1M 토큰당 USD)
IN_RATE, OUT_RATE = 0.09, 0.75

def cost(in_tokens: int, out_tokens: int) -> float:
    return in_tokens / 1_000_000 * IN_RATE + out_tokens / 1_000_000 * OUT_RATE

print(cost(1_000, 300))            # 챗봇 한 턴        -> $0.000315
print(cost(8_000, 500))            # RAG 답변 한 건    -> $0.001095
print(cost(500, 20) * 100_000)     # 10만 건 배치      -> ~$6.00

Kunavo pricing and Stripe billing

No subscriptions, no Google Cloud project. You top up a balance (using Stripe or local payment methods), and each call is deducted from it at the per-token rates above. It’s pay-as-you-go with a $10 minimum top-up, your balance never expires, and larger top-ups earn bonus credits. One balance covers Gemini and every other model — Claude, GPT, image, video, and audio — so you don’t have to reconcile a separate bill for each provider.

Which Gemini model should I choose?

  • gemini-2-5-flash — the default for chat, extraction, classification, summarization, and most RAG. Fast and the cheapest capable option.
  • gemini-3-1-pro — use it when Flash isn’t accurate enough: multi-step reasoning, code, vision, and very long context.

A good pattern is to route by difficulty: use Flash for common cases and upgrade to Pro only when a check fails. For a routing pattern in code, see the AI cost optimization guide.

How to reduce your Gemini bill

  1. Use a lower tier. Send the easy 80% to Flash; save Pro for the hard 20%.
  2. Limit output. Set max_tokens and stop sequences — output is the expensive side of the meter.
  3. Trim input. Retrieve fewer, better RAG chunks instead of stuffing your entire knowledge base into the context.
  4. Batch requests. Group independent calls to keep latency low and avoid retry storms.

Frequently asked questions

Is the Gemini API free?

Google AI Studio offers a rate-limited free tier for prototyping. For production, you pay per token. Kunavo is pay-as-you-go with a $10 minimum top-up — you pay per token at the rates below, your balance never expires, and no Google Cloud billing account is required.

How much does Gemini 2.5 Flash cost?

On Kunavo, Gemini 2.5 Flash costs $0.09 per 1M input tokens and $0.75 per 1M output tokens — about 70% below Google’s list price of $0.30 / $2.50. A typical chatbot turn (1K input, 300 output) costs about $0.0003.

Is Gemini cheaper than Claude or GPT?

Gemini 2.5 Flash is one of the cheapest capable models available — below Claude Haiku and most GPT tiers for high-volume workloads. For Claude’s rates and billing options, see Claude API pricing and billing explained; compare the full table on the pricing page.

How do I reduce Gemini API costs?

Drop down to Flash, limit output, trim retrieved context, and batch. Details are in the cost optimization guide. To start calling Gemini, see how to get a Gemini API key.