GPT API pricing is based on input and output rates per 1M tokens, with no monthly fee or free tier. This page compares the current rates for the GPT-5 and GPT-6 series with OpenAI's list prices, provides worked examples so you can calculate the cost of an actual request, explains two factors that can affect your bill (reasoning tokens and the long-input surcharge), and covers how to pay from South Korea.
Rates verified on October 4, 2026. OpenAI list prices are based on openai.com/api/pricing, and Kunavo rates are read live from the model catalog.
GPT-5 and GPT-6 Series Rate Table
| Model | OpenAI list price (input / output) | Kunavo (input / output) | Difference |
|---|---|---|---|
| GPT-6 Luna | $0.10 / $0.50 | $0.04 / $0.20 | Approximately 60% lower |
| GPT-6 Sol | $2.00 / $10.00 | $0.80 / $4.00 | Approximately 60% lower |
| GPT-5.6 Terra | $2.00 / $12.00 | $0.70 / $4.20 | Approximately 65% lower |
| GPT-5.5 | $5.00 / $30.00 | $2.00 / $12.00 | Approximately 60% lower |
| GPT-5.6 Sol | $5.00 / $30.00 Current promotional price: $4.00 / $20.00 (through at least November 21, 2026) | $2.00 / $12.00 | Approximately 60% lower approximately 40% below the promotional price |
| GPT-6 Astra | $10.00 / $50.00 | $4.00 / $20.00 | Approximately 60% lower |
The GPT-6 series is OpenAI's new generation. GPT-6 Astra launched as OpenAI's flagship model on September 3, 2026, followed by GPT-6 Sol and GPT-6 Luna on September 22, 2026. OpenAI says the two models build on the advances behind GPT-6 Astra, positioning Sol for complex coding and agent workflows, and Luna as “the most efficient model for focused, high-volume tasks.” Both models have a 1,050,000-token context window and a maximum output of 128,000 tokens.
GPT-6 Luna ($0.04 / $0.20) is the cheapest GPT on Kunavo. GPT-6 Sol ($0.80 / $4.00) has a lower output rate than GPT-5.6 Terra ($4.20). However, the input rate is lower for GPT-5.6 Terra ($0.70), so requests where input tokens exceed 2 times the output tokens are cheaper with Terra. If you send long prompts and receive short responses, calculate the cost for both models directly using the rates in the table above.
What does one actual request cost?
Rates alone can be hard to interpret, so calculate the cost yourself. The code below uses the same figures as the table above; change the token counts to estimate costs for your workload.
# Kunavo GPT 요율(1M 토큰당 USD): (input, output)
RATES = {
"gpt-5-6-terra": (0.70, 4.20),
"gpt-5-5": (2.00, 12.00),
"gpt-5-6-sol": (2.00, 12.00),
"gpt-6-sol": (0.80, 4.00),
"gpt-6-luna": (0.04, 0.20),
}
def cost(model: str, in_tokens: int, out_tokens: int) -> float:
i, o = RATES[model]
# 주의: GPT-5 계열 추론 모델의 reasoning 토큰은 출력 요율로 과금됩니다.
return in_tokens / 1_000_000 * i + out_tokens / 1_000_000 * o
print(cost("gpt-5-6-terra", 1_000, 300)) # 짧은 턴
print(cost("gpt-5-6-terra", 6_000, 500)) # RAG 답변
print(cost("gpt-5-5", 20_000, 3_000)) # 어려운 문제
print(cost("gpt-5-6-sol", 20_000, 3_000)) # GPT-5.6 플래그십
print(cost("gpt-6-sol", 20_000, 3_000)) # GPT-6 Sol
print(cost("gpt-6-luna", 1_000, 300)) # 가장 저렴한 짧은 턴Two factors that can affect your bill
1. Reasoning tokens are billed as output. GPT-5 series reasoning models generate internal reasoning tokens that are not visible on screen, and these tokens are charged at the output rate. Even a three-line answer may cost more than three lines' worth. For the most direct savings, limit max_tokens and lower the reasoning effort as far as the task allows.
2. Long-input surcharge. OpenAI applies 2× the input rate and 1.5× the output rate to the entire request, including cached input, when it exceeds 272,000 input tokens. Kunavo applies the same thresholds to GPT-5.5, GPT-5.6 Sol and Terra, and GPT-6 Astra, Sol, and Luna. Above the threshold, the rates in the table are multiplied by 2× for input and 1.5× for output, while the discount versus list price stays the same. For context-heavy tasks such as RAG over entire documents or codebase analysis, crossing this threshold can affect the bill more than model choice, so split inputs into chunks under 272,000 tokens where possible.
Connect your existing OpenAI code
Kunavo has an OpenAI-compatible endpoint, so you do not need to change SDKs. Just replace base_url and your key.
from openai import OpenAI
client = OpenAI(
api_key="sk-kn-...",
base_url="https://api.kunavo.com/v1", # 이 한 줄만 바꾸면 됩니다
)
resp = client.chat.completions.create(
model="gpt-5-6-terra",
messages=[{"role": "user", "content": "안녕하세요"}],
)
print(resp.choices[0].message.content)The same key works with Claude, image, and video models too — see the model catalog for the model list and Quickstart for endpoint details.
Payment in South Korea
Payments are processed through Stripe and are prepaid. Domestic and international credit and debit cards (Visa, Mastercard, Amex, JCB, UnionPay), Apple Pay, and Google Pay are supported; top up from a minimum of $10, charges are deducted according to usage, and the balance never expires. No OpenAI account or separate contract is required.
When the checkout is displayed in KRW, Kakao Pay, Naver Pay, PAYCO, Samsung Pay, and domestic cards are available as Kunavo top-up methods. Domestic cards include cards without international payments enabled, and Toss is not supported. These methods top up your Kunavo balance; they are not payments made directly to OpenAI. Prices are set in USD, and Stripe converts them to KRW at the exchange rate at the time of payment; this rate includes a 2–4% conversion fee paid by the purchaser. The fee does not apply when paying in dollars, but the domestic methods above appear only for KRW payments. See Billing documentation for all payment methods, and Claude API pricing and billing for the Claude cost overview.
Four ways to reduce costs
- Reduce output. Output is the most expensive part, including reasoning tokens. Limiting max_tokens and adjusting reasoning effort are the quickest ways to save.
- Choose a lower tier suited to the task. For high-volume tasks such as classification, extraction, and summarization, start with the least expensive GPT-6 Luna and move up only if needed. Luna's output rate is 1/100 of that of the top-tier GPT-6 Astra.
- Use caching for repeated context at the beginning of prompts (how it works).
- Batch tasks that do not need to run in real time. For tasks that do not make users wait, run them concurrently to increase throughput.
If you cannot pay for the OpenAI API, see How to pay for the OpenAI API.
Frequently asked questions
How much does the GPT API cost?
Pricing is per 1M tokens. October 4, 2026 According to OpenAI's official prices, GPT-5.6 Terra costs $2.00 input / $12.00 output; on Kunavo, the same model costs $0.70 / $4.20. Other Kunavo rates are GPT-6 Luna $0.04 / $0.20, GPT-6 Sol $0.80 / $4.00, GPT-5.5 $2.00 / $12.00, GPT-5.6 Sol $2.00 / $12.00, and GPT-6 Astra $4.00 / $20.00. GPT-6 Luna is the cheapest GPT among them.
How much does the GPT-6 API cost?
OpenAI's list prices (input/output per 1M tokens) for GPT-6 Sol and GPT-6 Luna, released on September 22, 2026, and GPT-6 Astra, released on September 3, 2026: GPT-6 Sol $2.00 / $10.00, GPT-6 Luna $0.10 / $0.50, GPT-6 Astra $10.00 / $50.00. Kunavo's input/output rates per 1M tokens: GPT-6 Sol $0.80 / $4.00, GPT-6 Luna $0.04 / $0.20, GPT-6 Astra $4.00 / $20.00. All three models are approximately 60% cheaper than list price. GPT-6 Luna is the cheapest GPT on Kunavo, and GPT-6 Sol has a lower output price than GPT-5.6 Terra ($4.20) (the input rate is lower for Terra: $0.70). For requests with more than 272,000 input tokens, the entire request is charged at 2 times the input rate and 1.5 times the output rate, exactly as with OpenAI.
Can I use the GPT API for free?
OpenAI does not offer a free tier for the GPT API — charges accrue per token from the first call. Kunavo is not free either, but it is pay-as-you-go rather than a subscription: add at least $10 and charges are deducted as you use the service, with no expiration on your balance. Months with no usage cost $0.
Are GPT-5 reasoning tokens billed separately?
GPT-5 series reasoning models generate reasoning tokens that are not visible on screen, and these tokens are billed at the output rate. That means your bill can be high even when the response looks short. Limit max_tokens and lower the reasoning effort as far as the task allows — output is always the more expensive side of the pricing table.
Do requests with long inputs cost more?
Yes. For requests with more than 272,000 input tokens, OpenAI applies 2× the input rate and 1.5× the output rate to the entire request, including cached input. Kunavo applies the same thresholds to GPT-5.5, GPT-5.6 Sol and Terra, and GPT-6 Astra, Sol, and Luna. The discount versus list price stays the same before and after the threshold. For context-heavy tasks such as RAG over entire documents or codebase analysis, crossing this threshold can affect the bill more than model choice.
How can I reduce GPT API costs?
Here are four ways, in order of impact: reduce output (including reasoning tokens, output is the most expensive), choose a lower tier suited to the task (for high-volume tasks such as classification and extraction, start with the least expensive GPT-6 Luna), use prompt caching for repeated context at the beginning of prompts, and batch tasks that do not need to run in real time. Switching base_url to a cheaper endpoint multiplies the savings.
How do I pay for GPT API usage from South Korea?
Kunavo processes payments through Stripe—it supports domestic and international credit and debit cards (Visa, Mastercard, Amex, JCB, UnionPay), Apple Pay, and Google Pay. It is prepaid, pay-as-you-go billing: top up from a minimum of $10 and charges are deducted according to usage; the balance never expires. If the checkout is displayed in KRW, Kakao Pay, Naver Pay, PAYCO, Samsung Pay, and domestic cards (including cards without international payments enabled) are also available. Toss is not supported. These are all methods for topping up your Kunavo balance, not payments made directly to OpenAI. No OpenAI account or separate contract is required. To get started, sign up and then issue a key.
Can I keep using my existing OpenAI SDK code as is?
Yes. Kunavo has an OpenAI-compatible endpoint, so change base_url to https://api.kunavo.com/v1 and replace the key. The rest of your code can stay as it is, and the same key works with Claude, image, and video models too.