Token counter
Free, instant, no signup. Paste any text — a prompt, a system message, a document, a chat log — and get its token count, its characters and words, and what that text costs as input or as output on every Claude and GPT model, each next to its provider's list price. Counting runs in your browser.
- tokens
- 0
- characters
- 0
- words
- 0
Counted in your browser. The tokenizer downloads once, the first time you enter text, and nothing you paste is sent anywhere.
What this text costs, per request
Paste text above to see how many tokens it is and what it costs on each model.
Exact for GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.5: gpt-tokenizer maps these models to o200k_base, the tokenizer counted here.
≈ marks an estimate: the o200k_base count of the same text. GPT-6 is newer than the tokenizer's model list, so its tokenizer is not confirmed there, and Claude tokenizes with its own vocabulary.
Anthropic does not publish Claude's tokenizer. Its count_tokens endpoint is free and returns Claude's own number, and Anthropic's docs say Claude 4.7 and later models produce about 30% more tokens than earlier Claude models for the same text.
The same text takes more tokens in other languages
Kunavo's guides are published in thirteen languages, so one piece of content can be counted thirteen times. Across 266 guides — 13,501 passages and 586,465 English words — this is how many o200k_base tokens each language needed, with English as 1:
| Language | Tokens vs English | Characters per token |
|---|---|---|
| English | 1.00× | 4.62 |
| Chinese (Simplified) | 1.18× | 1.76 |
| Indonesian | 1.21× | 4.41 |
| Portuguese (Brazil) | 1.22× | 4.38 |
| Spanish | 1.23× | 4.52 |
| German | 1.29× | 4.39 |
| French | 1.31× | 4.41 |
| Italian | 1.35× | 4.07 |
| Vietnamese | 1.38× | 3.72 |
| Chinese (Traditional) | 1.39× | 1.50 |
| Korean | 1.41× | 2.05 |
| Turkish | 1.44× | 3.62 |
| Japanese | 1.64× | 1.64 |
English prose came to 1.30 tokens per word, so 1,000 words is about 1,300 tokens and a million tokens is about 770,000 words. Priced on English prompts, the same product costs about 64% more in input tokens in Japanese, and about 41% more in Korean.
Two limits on reading it. The translations are Kunavo's own, so a tighter or looser translation would move a row. And these are OpenAI's o200k_base counts: Claude's tokenizer splits text its own way, so for Claude the multiples will not be the same. To turn a count into a monthly bill, take it to the AI token cost calculator.
Price what you counted
Price a workload per month
Cache hits and writes priced
Cached input and reasoning
Every Claude rate, vs list
GPT rates by model
Every model, vs official
How many tokens is 1,000 words?
About 1,300 tokens for English prose with o200k_base, the tokenizer of OpenAI's GPT-4o and GPT-5 models: across 586,465 words of Kunavo's own guides the average was 1.30 tokens per word. Code, numbers, URLs and most other languages take more tokens per word, so paste your own text above for the real figure.
How many words is 1 million tokens?
Roughly 770,000 English words, at the 1.30 tokens per word measured on Kunavo's guides with o200k_base. A million tokens holds fewer words in other languages: the same guides took 1.64 times as many tokens in Japanese, 1.29 times in German and 1.18 times in Simplified Chinese.
Is the token count exact for Claude?
No — for Claude it is an estimate. Anthropic does not publish Claude's tokenizer, so the figure shown for Claude models is the o200k_base count of the same text. Anthropic's token counting endpoint, /v1/messages/count_tokens, is free and returns Claude's own number, and Anthropic's documentation says Claude 4.7 and later models produce about 30% more tokens than earlier Claude models for the same text.
Which models is the count exact for?
GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.5: gpt-tokenizer, the open-source library that does the counting here, maps those models to o200k_base, the encoding it counts with. GPT-6 is newer than that library's model list, so its count is marked as an estimate rather than assumed. Either way the figure is the text itself — a chat request adds a few formatting tokens per message.
Is the text I paste uploaded anywhere?
No. The tokenizer runs in your browser: its vocabulary downloads once, the first time you enter text, and after that every count happens on your device. The text is not sent to Kunavo or to anyone else, and it is gone when you close the page.
Why does my API usage show more input tokens than this counter?
Because a request carries more than your text. Chat formatting adds a few tokens per message, the system prompt and tool definitions are sent on every call, images are billed as tokens of their own, and reasoning models bill hidden reasoning as output. The usage field of each API response is the number you are billed for; this counter is for sizing a prompt before you send it.