Back to guides
Cost·September 11, 2026·Updated October 5, 2026·9 min read

Gemini API Pricing 2026 — Free Tier, Model Prices, and Gemini 3.8 Flash

“Gemini pricing” is really three things: the free tier, app subscriptions, and API token-based billing. This page separates them, lists current model prices, and honestly identifies which model is cheaper to buy directly from Google.

Gemini API costs are billed by token, with no monthly fee. To try it for free, Google AI Studio has a rate-limited free tier; for production, you pay per 1M tokens, with separate rates for input and output. As of October 5, 2026, Kunavo’s Gemini 3.8 Flash costs $0.525 for input / $2.625 for output — Google currently charges a limited-time promotional price of $0.75 / $3.75 for its 3.x Flash models (through December 31, 2026; from 2027, the standard price returns to $1.50 / $7.50), and Kunavo is about 30% below that promotional price. The previous-generation Gemini 2.5 Flash costs $0.09 / $0.75, and Gemini 3.1 Pro costs $0.70 / $4.20.

There’s one exception to mention first: Gemini 3.6 Flash. On Kunavo it costs $1.05 / $5.25, approximately 40% higher than Google's discounted price — during the promotional period, calling Google directly is cheaper. 3.7 and 3.8 Flash are both newer and cheaper than it, So don’t choose 3.6 based on price.

Rates verified on October 5, 2026. See ai.google.dev/gemini-api/docs/pricing for Google’s official prices and free-tier details; Kunavo’s rates are shown in our model catalog and match the amounts billed by the API.

“Gemini cost” actually means three different things

RouteWhat you pay forWhether you get an API keyBest for
Free tier (Google AI Studio)$0, rate-limitedYes (limited)Prototyping, learning, personal small tools
Google AI subscription (Pro / Ultra in the Gemini app)Fixed monthly feeNoUse only in the Gemini app; no coding
API pay-as-you-goPer token, no monthly feeYesIntegrate into a product, process batches, control costs

The most common misconception is in the middle row: a Google AI Pro subscription does not give you more API credits. The app’s monthly fee and the developer API are billed separately. To make calls from code, you still need to use the free tier or pay by token. Conversely, if you only chat in the Gemini app and don’t write code, the free version or a subscription is enough; you don’t need the API.

Is the Gemini API free? What can you do with the free tier?

There is a free tier, and it really costs $0: create a key in Google AI Studio and calls within the rate limits are free. It’s more than enough for trying models and building prototypes. Google adjusts the specific limits, so this page doesn’t repeat figures that may become outdated — see Google’s pricing page and rate limits documentation.

Its limits are also clear: it can’t handle production traffic, and the figures can change at any time, so you need to pay for production. Kunavo has no free tier, but also has no monthly fee — the minimum prepaid balance is $10, you pay for what you use, and months with no calls cost $0.

Gemini API prices — per 1M tokens by model

ModelKunavo inputKunavo outputGoogle official (input / output)Compared with Google’s current prices
Gemini 3.8 Flash$0.525$2.625Promotional price $0.75 / $3.75 (from 2027: $1.50 / $7.50)approximately 30% lower than Google's discounted price
Gemini 3.7 Flash$0.525$2.625Promotional price $0.75 / $3.75 (from 2027: $1.50 / $7.50)approximately 30% lower than Google's discounted price
Gemini 3.6 Flash$1.05$5.25Promotional price $0.75 / $3.75 (from 2027: $1.50 / $7.50)approximately 40% higher than Google's discounted price
Gemini 3.1 Pro$0.70$4.20$2.00 / $12.00Save approximately 65%
Gemini 2.5 Flash$0.09$0.75$0.30 / $2.50Save approximately 70%

Those three rows for 3.x Flash compare with the promotional prices Google is actually charging today, not its standard prices for 2027. Compared with the standard prices, the figures would look much better, but those aren’t what you’re paying Google now. Gemini 3.1 Pro has another pricing rule: when a single request’s prompt exceeds 200K tokens, the entire request is billed at 2× the input rate / 1.5× the output rate. Specifications for each model are on the Gemini 3.8 Flash, Gemini 3.1 Pro, and Gemini 2.5 Flash model pages; all rates are on the pricing page.

Gemini 3.8 Flash pricing — what one request actually costs

A rate table is less useful than seeing the cost of one request. The examples below are all calculated from the rates above:

Use caseTokens (input / output)3.8 Flash (Kunavo)3.8 Flash (Google promotional price)2.5 Flash3.1 Pro
One chatbot turn1,000 / 300$0.0013$0.0019$0.00031$0.002
One RAG answer8,000 / 500$0.0055$0.0079$0.0011$0.0077
One long-document analysis20,000 / 2,000$0.016$0.022$0.0033$0.022
Classifying 100,000 documents500 / 20 × 100,000$31.50$45.00$6.00$43.40

One thing that can push up your bill: thinking tokens. Gemini 3.x Flash and 3.1 Pro both “think” before answering. Thinking tokens don’t appear in the response, but are billed at the output rate. They aren’t included in the table above, so difficult tasks will cost more than the table shows — use usage for the number of output tokens in the response. Calculate using your own numbers:

gemini_cost.py
# Gemini API 的費用,直接用實際 token 數算。
# 費率為每 1M token USD:(input, output)
RATES = {
    "gemini-3-8-flash": (0.525, 2.625),
    "gemini-3-1-pro":   (0.70, 4.20),
    "gemini-2-5-flash": (0.09, 0.75),
}

def cost(model, inp, out):
    i, o = RATES[model]
    return inp / 1e6 * i + out / 1e6 * o  # 思考 token 按輸出費率計

print(cost("gemini-3-8-flash", 8_000, 500))         # RAG 回答一次
print(cost("gemini-2-5-flash", 500, 20) * 100_000)  # 10 萬筆分類

Call Gemini with your existing OpenAI SDK

Kunavo is an OpenAI-compatible endpoint: replace base_url and the key, set model to the Gemini model, and you’re ready to go. You don’t need a Google Cloud project or billing account. The same key can also call Claude, GPT, and image and video models.

gemini_kunavo.py
from openai import OpenAI

client = OpenAI(
    api_key="sk-kn-...",
    base_url="https://api.kunavo.com/v1",  # 只改這一行
)

resp = client.chat.completions.create(
    model="gemini-3-8-flash",
    messages=[{"role": "user", "content": "用三句話說明什麼是 token"}],
)
print(resp.choices[0].message.content)
print(resp.usage)  # 這次請求計費用的 token 數

For the full steps, see Quick start. To get a key directly from Google, see How to get a Gemini API key (English).

Honestly: when it’s better to go directly to Google

  • You need guaranteed quota or an SLA. The gateway uses shared capacity: there is no dedicated quota or contractually guaranteed SLA. If you’re at a scale that needs either, apply directly to Google.
  • You can only use Gemini 3.6 Flash. During the promotional period, Google charges less than Kunavo.
  • You only chat in the app. If you don’t write code, the free version or a Google AI subscription is enough; you don’t need the API.

This route is for using one key to call Gemini, Claude, and GPT, avoiding Google Cloud billing, or keeping costs in check with a prepaid balance.

Paying from Taiwan

Paying Google directly requires enabling billing for the project in Google Cloud and linking an international credit card. Kunavo balance uses Stripe and supports cards including JCB (Visa, Mastercard, American Express, JCB, UnionPay), Apple Pay, and Google Pay; it is prepaid, with a minimum top-up of $10 and no balance expiration, and failed requests are not charged. Top-ups of $100 or more receive an additional bonus ($100 credited as $110).

Taiwan has no local payment options — JKoPay and LINE Pay are not on the supported list, so an international card is the only option. See the billing documentation for details.

How to reduce Gemini costs

  1. First, choose the right model. If you want a 3.x generation model, use 3.8 or 3.7 Flash: they’re newer and cheaper than 3.6; Assign large volumes of simple tasks such as classification and extraction to 2.5 Flash.
  2. Ask the model to write less. Output costs more than input, so specifying the output format and length in your prompt has the most direct effect.
  3. Include less context. Retrieve a few relevant RAG chunks instead of sending the entire knowledge base — the full input is billed every time.
  4. Keep 3.1 Pro prompts at or below 200K tokens. Above that, the entire request is billed at 2× the input rate / 1.5× the output rate.

For routing patterns, see the AI cost optimization guide (English). To compare other providers using the same method, see OpenAI API pricing and Claude pricing.

Frequently asked questions

How much does the Gemini API cost?

The Gemini API is billed by token, with no monthly fee and separate rates for input and output. As of October 5, 2026, Kunavo’s rates per 1M tokens are Gemini 3.8 Flash: $0.525 input / $2.625 output; Gemini 3.1 Pro: $0.70 / $4.20; Gemini 2.5 Flash: $0.09 / $0.75. Google currently offers a limited-time promotional price for its 3.x Flash models at $0.75 / $3.75 (through December 31, 2026; from 2027, the price is $1.50 / $7.50). Google AI Studio also has a rate-limited free tier for trying it out.

Can I use the Gemini API for free?

You can try it, but the free tier is not suitable for production. Google AI Studio offers a rate-limited free tier, with calls within the limits costing $0; Google changes the limits, so check its Gemini API pricing page for the current figures. Production traffic requires payment: pay Google directly or use a gateway billed by token. Kunavo has no free tier, but also has no monthly fee — the minimum prepaid balance is $10, you pay for what you use, and months with no calls cost $0.

How much does Gemini 3.8 Flash cost?

As of October 5, 2026, Kunavo’s Gemini 3.8 Flash rates are $0.525 per 1M input tokens / $2.625 per 1M output tokens. Google currently offers a limited-time promotional price for its 3.x Flash models at $0.75 / $3.75 through December 31, 2026. Kunavo’s price comparison: approximately 30% lower than Google's discounted price. On January 1, 2027, Google returns to the standard price of $1.50 / $7.50. One RAG answer (8,000 input / 500 output tokens) costs about $0.0055 on Kunavo.

Is Gemini 3.6 Flash still worth using?

It is not recommended to choose it based on price. Kunavo's Gemini 3.6 Flash is $1.05 / $5.25, approximately 40% higher than Google's discounted price, so during the promotional period, calling 3.6 directly through Google is actually cheaper. 3.7 Flash and 3.8 Flash are both newer than 3.6, and at Kunavo they are also cheaper than 3.6 ($0.525 / $2.625). Unless your code specifically requires 3.6, switch to a newer model.

I have Google AI Pro. Can I use the Gemini API?

No. Paid subscriptions to the Gemini app (Google AI Pro / Ultra) and the developer Gemini API are separate products with separate billing. The monthly fee does not include API credits or an API key. To make calls from code, you still need to use the free tier or pay by token; conversely, having an API key does not add features to the app.

Can I use the Gemini API without a Google Cloud billing account?

You can make calls through an OpenAI-compatible gateway. With Kunavo, one key beginning with sk-kn- can call Gemini; no Google Cloud project or billing account is required, and Claude and GPT use the same balance. The tradeoff is that this route uses shared capacity: there is no dedicated quota or contractually guaranteed SLA. For a scale that requires guaranteed quota or an SLA, apply directly to Google.

How do I pay for the Gemini API in Taiwan?

Paying Google directly requires enabling billing for the project in Google Cloud and linking an international credit card. Kunavo uses Stripe prepaid billing: it supports Visa, Mastercard, American Express, JCB, UnionPay, Apple Pay, and Google Pay, with a minimum top-up of $10, no balance expiration, and no charge for failed requests. Taiwan has no local payment channels—JKO Pay and LINE Pay are not on the available list.