Independent Research on Mutual Funds, Stocks & IPOs for Indian Investors

top of page

What Are Tokens in Claude, OpenAI, and Gemini?

  • 14 minutes ago
  • 5 min read

If you've used an AI chatbot for more than a few minutes, you've probably run into the word “token”, in a pricing page, an error message about hitting a limit, or a note about a model's “context window.” Tokens are the basic unit every large language model actually reads and writes in.


Understanding them helps explain why a long conversation suddenly gets cut off, why the same prompt costs different amounts on different models, and why a document that looks short to you can quietly eat up a huge chunk of a model's memory.


A token is a chunk of text, sometimes a whole word, sometimes just part of one, sometimes a single punctuation mark or space. Before a model can process anything you type, it breaks your text into these chunks and converts each one into a number. The model then works entirely in numbers; converting back into readable words only happens at the very end, when it generates a response.


A useful rule of thumb, and one that Claude, OpenAI, and Google's own documentation all converge on: roughly 1 token equals about 4 characters of English text, or about 0.75 words. So 100 tokens is roughly 60 to 80 English words. It's an approximation, not an exact conversion; code, non-English languages, and unusual formatting all tokenize differently.


Why Tokens Matter

Three things in particular hinge on tokens:

• Context window, the maximum number of tokens a model can hold in a single conversation, counting your input, any files or history you've included, and the model's own reply.

• Cost, since API usage is billed per token, with input and output typically priced differently (output usually costs several times more per token than input).

• Speed, since more tokens generally means more computation, which shows up as latency.


None of this is unique to one company. Every major model works this way. What differs is how each one breaks text into tokens in the first place, and that's where Claude, OpenAI, and Gemini actually diverge.


How Claude Tokenizes Text

Claude uses a proprietary tokenizer built on byte pair encoding (BPE), a method that starts from individual characters and progressively merges the most frequently occurring pairs into single tokens. It's trained on Anthropic's own data, so while the underlying technique is similar in spirit to what OpenAI uses, the specific vocabulary is different, which means the same sentence will not produce identical token counts on Claude and GPT.


Most current mainstream Claude models work with a context window of around 200,000 tokens. Anthropic's API includes a free token counting endpoint, so you can check exactly how many tokens a prompt will use before sending it, rather than relying on the character or word rule of thumb.


One detail worth knowing: Claude's tokenizer isn't fixed even within Anthropic's own model lineup. The tokenizer introduced with Claude Opus 4.7 produces roughly 30% more tokens for the same piece of text than the tokenizer used in earlier Claude models, a difference carried forward into Claude Fable 5 and Claude Mythos 5.


If you're estimating cost or context usage across Claude model generations, a token count measured on one model doesn't automatically transfer to another, it's worth recounting with the specific model you intend to use.


How OpenAI Tokenizes Text

OpenAI's GPT models use byte pair encoding as well, implemented through an open source library called tiktoken, which you can run yourself for free. Different model generations use different “encodings,” essentially, different trained vocabularies.


Newer models generally use an encoding with a vocabulary of around 200,000 tokens, while the GPT-3.5 and GPT-4 generation used a smaller one of around 100,000, and older GPT-3 and Codex era models used encodings smaller still. Because tiktoken is open source, it's the one tokenizer among the three you can install locally and get an exact, not estimated, count.


Context windows vary considerably across OpenAI's model lineup, from roughly 128,000 tokens on some models up to around 1,000,000 on others. Because this changes with nearly every model release, it's worth checking OpenAI's current model documentation rather than assuming a figure carries over from an older model you've used before.


How Gemini Tokenizes Text

Gemini takes a genuinely different approach. Instead of byte pair encoding, Google uses SentencePiece with a Unigram language model. The practical difference: BPE builds tokens by greedily merging the most common adjacent pairs, while a Unigram model considers many possible ways a piece of text could be split and picks the statistically most probable segmentation.


Gemini's vocabulary sits at roughly 256,000 tokens.

Gemini is also the provider most associated with very large context windows. Depending on the specific model and how you're accessing it, context windows commonly reach 1,000,000 tokens, with some models and tiers extending to 2,000,000.


As with the other two providers, exact figures vary by model and change frequently enough that Google's own model documentation is the place to confirm a specific number rather than a general figure like the one above.


Why The Same Text Produces Different Token Counts

Because each company trains its own tokenizer on its own data, with its own vocabulary size and its own algorithm, there's no universal token. A sentence that comes out to 40 tokens on one model might come out to 35 or 48 on another. A word like “tokenization” might be a single token in one vocabulary and split into two or three pieces in another. This is normal and expected, not a bug in any one system.


The practical consequence is that token based comparisons across providers are never perfectly apples to apples. A context window advertised at 1,000,000 tokens on one provider doesn't necessarily hold exactly the same amount of English prose as 1,000,000 tokens somewhere else, though the roughly 4-characters-per-token heuristic holds up reasonably well across all three for plain English text. It tends to break down faster for code, tables, non-English languages, and anything with unusual formatting.


A Quick Comparison

 

Claude

OpenAI (GPT models)

Gemini

Tokenizer approach

Proprietary byte pair encoding

Byte pair encoding, via the open source tiktoken library

SentencePiece with a Unigram language model

Approximate vocabulary

Not publicly disclosed

Roughly 100,000 to 200,000, depending on model generation

Roughly 256,000

Rough rule of thumb

About 4 characters or 0.75 words per token

About 4 characters or 0.75 words per token

About 4 characters per token

Typical context window

Commonly around 200,000 tokens on current mainstream models

Varies widely by model, roughly 128,000 up to around 1,000,000

Commonly around 1,000,000, with some models and tiers reaching 2,000,000

Free token counting

API token counting endpoint

Open source tiktoken library, or OpenAI's online tokenizer tool

API count_tokens call

Figures above are current as of mid-2026 and change with nearly every model release. Treat them as a general sense of scale rather than a number to build cost estimates on, and check each provider's own documentation for the specific model you're using.


A Few Practical Takeaways

• Count before you send, especially for large prompts. All three providers offer a free way to check token counts ahead of time. It's more reliable than estimating from word or character counts, particularly for anything other than plain English prose.

• Don't carry a token count between models. Switching from one provider to another, or even between model generations within the same provider, can change the count for identical text. Recount rather than assume.

• Leave room for the output. A context window covers your input and the model's reply together. A prompt that uses nearly all of a model's window leaves little space for a detailed answer.


  • X
  • LinkedIn
  • Instagram
  • Facebook

Warning: Investment in Mutual Funds and  Securities Market are subject to market risks. Read all scheme related documents carefully before investing.

Disclaimer: This website provides educational content only and does not offer investment advice.

List of mutual fund companies (AMCs):  ONE  |  Abakkus  |  Aditya Birla Sun Life  |  Angel One  |  Axis  |  Bajaj Finserv  |  Bandhan  |  Bank of India  |  Baroda  |   BNP Paribas  |  Canara Robeco  |  Capitalmind  |  Choice  |  DSP  |  Edelweiss  |  Franklin Templeton  |  Groww  |  HDFC  |  Helios  |  HSBC  |  ICICI Prudential  | Invesco  |  ITI  |  JioBlackRock  |  JM Financial  |  Kotak Mahindra  |  LIC  |  Mahindra Manulife  |  Mirae Asset  |  Motilal Oswal  |  Navi  |  Nippon India  |  NJ  |  Old Bridge  |  PGIM India  |  PPFAS  |  Quant  |  Quantum  |  Samco  |  SBI  |  Shriram  |  Sundaram  |  Tata  |  Taurus  |  The Wealth Company  |  TRUST  |  Unifi  |  Union  |  UTI  |  WhiteOak  |   Capital  |  Zerodha

© 2026 by Equity Research India

bottom of page