What it means
A token is a small chunk of text, such as a short word, part of a longer word or a punctuation mark, that an AI model reads and writes.
Large language models don’t read text letter by letter or word by word. A piece of software called a tokeniser first chops your message into tokens. Common words like “the” or “house” usually become one token each. Rarer words get split, so “tokenisation” might come out as something like “token”, “is” and “ation”. The exact split varies by model, and spaces, punctuation and emoji count too.
Picture a fridge magnet poetry kit. It has magnets for common whole words, plus fragments like “un”, “ing” and “ly” for building everything else. A tokeniser works from a similar fixed set of pieces, and the model writes its reply by adding one piece at a time.
The handy rule for English is that a token is about three quarters of a word. So 100 tokens is roughly 75 words. It’s an average, not a law. Code, numbers, unusual names and most other languages use more tokens for the same meaning.
Where it came from
“Token” is an old word for a sign or a stand-in, as in a bus token or a token of thanks. Computing borrowed it decades ago. A compiler’s first job is to break code into tokens such as names, numbers and symbols. In security, a token is a digital key that proves who you are. Crypto later added a meaning of its own.
Many AI tokenisers use a method called byte-pair encoding, which began as a data compression technique. It repeatedly merges the pairs of characters or chunks that appear together most often in a large body of text. That text is usually dominated by English, so English words tend to become whole tokens. Many other languages, especially those in non-Latin scripts such as Hindi or Thai, get chopped into smaller fragments, so the same sentence can need several times as many tokens.
The pricing sense is newer. The usual account is that it took hold around 2020, when companies began selling access to language models through APIs. Tokens were the obvious unit to charge by, because they’re what the model actually processes. Prices are now usually quoted per million tokens, with separate rates for input (what you send) and output (what the model writes). Output typically costs more.
Limits work the same way. A model’s context window is the most tokens it can handle at once, counting your message, attached files, the conversation so far and its reply. Early chat models managed a few thousand. Many current ones handle hundreds of thousands, and some pass a million.
Here are rough sizes for ordinary English text.
| Text | Rough size in tokens |
|---|---|
| A short, common word | 1 |
| A long or unusual word | Often 2 to 4 |
| A 75-word paragraph | About 100 |
| A 750-word article | About 1,000 |
| A 90,000-word novel | About 120,000 |
How people actually use it
- “How many tokens is that prompt?” A developer checking cost, or whether the text will fit. Trimming prompts is a routine part of prompt engineering.
- “We hit the token limit halfway through the contract.” The document was too long for the model’s context window.
- “Input’s cheap. It’s the output tokens that add up.” Work chat about an AI bill.
- “This agent burns through tokens.” A complaint that an AI agent, which works through many steps on its own, runs up a big bill.
- “It’s doing about 80 tokens a second.” Speed talk. Tokens per second measures how fast a model writes.
In a sentence
Kai: why did the AI bill triple
Rosa: someone left the agent running all weekend. millions of tokens
Lena: it forgot what I told it at the start
Omar: long chat? you probably went past the context window
Lena: in English please
Omar: too many tokens. it ran out of room
Ines: same prompt in Hindi costs more. is that a bug
Dev: no, Hindi just splits into more tokens
Common misconceptions
- A token isn’t a word. It’s often a piece of one. On average, English runs at about four tokens for every three words.
- Tokens aren’t only about money. They also set how much a model can read at once and how long it takes to reply.
- Models don’t see individual letters. Because they work in chunks, models can stumble on tasks like counting the letters in a word.
Related terms
- Tokeniser: the software that splits text into tokens, spelled “tokenizer” in most documentation.
- Context window: the most tokens a model can handle at once, input and output combined.
- Input and output tokens: what you send and what the model sends back, usually priced separately.
- Tokens per second: a common measure of how fast a model produces text.
- Embedding: the list of numbers a model turns each token into before working with it.
Questions people ask
How many words is 1,000 tokens?
In English, 1,000 tokens is roughly 750 words. The exact figure depends on the text, and code or other languages usually fit fewer words into the same number of tokens.
Why do AI companies charge by the token?
AI companies charge by the token because tokens are what their models actually process, so the count tracks the computing work closely. Longer prompts and answers cost more to run.
Why do some languages use more tokens than English?
Tokenisers are mostly built from English-heavy text, so English words often become single tokens. Many other languages get split into smaller pieces, which means more tokens, higher costs and less room in the context window.
Are AI tokens the same as crypto tokens?
AI tokens and crypto tokens are unrelated. An AI token is a chunk of text a model reads or writes, while a crypto token is a digital asset recorded on a blockchain.
The short version
A token is the chunk of text an AI model reads and writes, usually a short word or part of a longer one. In English, a token is about three quarters of a word. Tokens decide what you pay, how much a model can read at once and how fast it replies.