Published on
Have you ever seen «token limit» or «price per token» and pretended to understand? No worries: the word sounds technical, but the idea is very simple.
In this lesson you will find out what a token is, how artificial intelligence uses tokens to read your question and to write its answer one small piece at a time, and why it is worth knowing. There are four examples where you watch an answer being born token by token.
Download the handoutPrintable PDF version of this lesson, with exercise and glossary.
Download the handout as PDFPlay with tokens: Tokenize MePick a question and watch how AI splits it into tokens and writes the answer one piece at a time. A free interactive game with animations and sound effects.
Play nowUnderstand what a token is, how AI turns text into tokens and builds an answer by adding one at a time, so you can use chatbots better and make sense of limits and costs.
You see words and sentences. Inside, the computer sees only numbers.
When you write «Hello, how are you?» to a chatbot, nothing inside the machine «understands» those letters the way you do. First the text has to be taken apart into small pieces, and each piece turned into a number. Only then can the model work with it.
Those pieces are called tokens. Think of toy building bricks: with just a few kinds of brick you can build castles, spaceships and even a platypus. With a limited set of tokens, AI builds any sentence.

A token is a small piece of text: the unit AI uses to read what you write and to write what it answers.
A token is not always a word. It can be a whole word, part of a word, a punctuation mark, or a space attached to the word that follows. Common words, which AI has seen millions of times, are usually a single token. Rare words are split into several pieces.
| Text | How it might be split | Tokens |
|---|---|---|
| cat | «cat» | 1 |
| Hello, how are you? | «Hello» «,» «·how» «·are» «·you» «?» | 6 |
| platypus | «pl» «aty» «pus» | 3 |
A handy rule of thumb, especially for English: one token is about four characters, and one hundred tokens are about seventy-five words. Languages with long words and many endings, like Italian, usually use somewhat more tokens than the same text in English.

These are estimates. Every model has its own way of splitting text, so the same text can have a different token count from one service to another. For an exact number, use the counter of the service you are using.
AI has a closed list of tokens it knows, the vocabulary: tens of thousands of pieces, each with a code number. The text you write is cut into pieces from that list, and each piece is replaced by its number. That series of numbers is what goes into the model.

The same thing happens in reverse: the model produces a number, and the number is translated back into the piece of text you see appear on the screen.
AI does not write the answer all at once: it builds it one piece at a time, like when you compose a message on your phone.
The mechanism always repeats the same way. The model reads all the text so far, calculates how likely each possible token is to come next, picks one and adds it. Then it starts again, with one more token to read.

It is like the autocomplete on your phone keyboard, but trained on an enormous amount of text, and starting over after every single piece.
The model does not always pick the most likely token: sometimes it also picks among less likely ones. A setting called temperature controls how adventurous it is. Low: predictable, uniform answers. High: more imagination and more surprises. That is why asking the same thing twice gives you two different texts.
Let us see it live. Each example has a question and the answer that is born one token at a time. Read the tables from top to bottom: each row adds a piece, and the last column shows the text that exists at that point. The splits are illustrative, but the mechanism is the real one.
«What is the capital of Italy?»
| Step | Token added | Text built so far |
|---|---|---|
| 1 | «The» | The |
| 2 | «·capital» | The capital |
| 3 | «·of» | The capital of |
| 4 | «·Italy» | The capital of Italy |
| 5 | «·is» | The capital of Italy is |
| 6 | «·Rome» | The capital of Italy is Rome |
| 7 | «.» | The capital of Italy is Rome. |

At step 6 the model chooses among many candidates, each with its own probability. This is how that moment might look:
| Candidate for step 6 | Probability (illustrative) |
|---|---|
| «·Rome» | 97% |
| «·Milan» | 1% |
| «·Naples» | less than 1% |
| all other tokens | the rest |
«Give me one sentence of breakfast advice.»
| Step | Token added | Text built so far |
|---|---|---|
| 1 | «Start» | Start |
| 2 | «·with» | Start with |
| 3 | «·a» | Start with a |
| 4 | «·coffee» | Start with a coffee |
| 5 | «·and» | Start with a coffee and |
| 6 | «·a» | Start with a coffee and a |
| 7 | «·slice» | Start with a coffee and a slice |
| 8 | «·of» | Start with a coffee and a slice of |
| 9 | «·bread» | Start with a coffee and a slice of bread |
| 10 | «·and» | Start with a coffee and a slice of bread and |
| 11 | «·jam» | Start with a coffee and a slice of bread and jam |
| 12 | «.» | Start with a coffee and a slice of bread and jam. |

«Which Australian animal has a duck-like bill?»
| Step | Token added | Text built so far |
|---|---|---|
| 1 | «It» | It |
| 2 | «'s» | It's |
| 3 | «·the» | It's the |
| 4 | «·pl» | It's the pl |
| 5 | «aty» | It's the platy |
| 6 | «pus» | It's the platypus |
| 7 | «.» | It's the platypus. |

Common words cost one token, rare ones cost more. The model rebuilds them brick by brick.
«Write the silliest excuse for being late.»
| Step | Token added | Text built so far |
|---|---|---|
| 1 | «A» | A |
| 2 | «·seagull» | A seagull |
| 3 | «·stole» | A seagull stole |
| 4 | «·my» | A seagull stole my |
| 5 | «·breakfast» | A seagull stole my breakfast |
| 6 | «,» | A seagull stole my breakfast, |
| 7 | «·so» | A seagull stole my breakfast, so |
| 8 | «·I» | A seagull stole my breakfast, so I |
| 9 | «·chased» | A seagull stole my breakfast, so I chased |
| 10 | «·it» | A seagull stole my breakfast, so I chased it |
| 11 | «·to» | A seagull stole my breakfast, so I chased it to |
| 12 | «·the» | A seagull stole my breakfast, so I chased it to the |
| 13 | «·sea» | A seagull stole my breakfast, so I chased it to the sea |
| 14 | «.» | A seagull stole my breakfast, so I chased it to the sea. |

The funniest choice is at step 2. When the request is creative, no candidate dominates and temperature makes the difference:
| Candidate for step 2 | Probability (illustrative) |
|---|---|
| «·seagull» | 22% |
| «·pigeon» | 19% |
| «·penguin» | 8% |
| «·alien» | 4% |
| all other tokens | the rest |

If you think tokens are boring, you have not yet seen what happens when AI has to cut up certain words.
The longer, rarer or weirder a word is, the more pieces are needed to rebuild it. Here is our little museum: the cuts are examples, but the principle is real.

Which costs more: «hello» or «HELLOOOOO»? Guess, then check with the counter.
The cuts shown are illustrative: every model has its own cutter, and real results can differ.
Tokens are not an engineer's curiosity: they explain three things you meet every day when you use a chatbot.
Many AI services, especially those for businesses and developers, are paid by the token: both the ones you send and the ones you receive are counted. Short prompts and focused answers cost less. Prices change often: always check the service's current price list.
The model can only take into account a limited amount of text, the context window, measured in tokens. Your question, the conversation so far and the answer all fit inside it. When the conversation gets too long, the oldest parts fall out and the model «forgets» them.
The model sees pieces, not letters. That is why it can stumble on tasks that are trivial for us, like counting how many times a letter appears in a word or spelling a word backwards. It is not stupid: it is looking at the bricks, not at the individual grains of plastic.
| Situation | What happens | What to do |
|---|---|---|
| Very long text | You hit the limits or pay more | Summarize, split into parts, cut the fluff |
| Never-ending conversation | The model loses early details | Start over and paste a short summary |
| Question about single letters | The model sees pieces, not letters | Always check the result yourself |
| Text in Italian | Usually costs a few more tokens | Don't worry: it only matters at large volumes |
A token is a brick. AI has never written a whole sentence in one go: it has always stacked one piece on top of another.
— Gianluca Bonomo
Play with tokens: Tokenize MePick a question and watch how AI splits it into tokens and writes the answer one piece at a time. A free interactive game with animations and sound effects.
Play nowFifteen minutes, a browser and your favourite chatbot. Nothing to install.
Can you say how many tokens your sentence has?
Did you notice which words were split into several pieces?
Can you explain to a colleague why a phone keyboard looks like an AI?
No. Sometimes it matches a word, sometimes it is a piece of a word, a punctuation mark or a space. Common words are usually one token, rare ones use more.
It depends on the model and the language. As an order of magnitude, for English one hundred tokens are about seventy-five words; for other languages the count is usually higher. For an exact number use the service's counter.
Because at each step the model chooses among several likely tokens and does not always take the first one. Temperature controls how adventurous it is: the higher it is, the more the answers vary.
With many services for businesses and developers, yes: tokens sent and received are paid for, at rates that change. Consumer subscriptions usually apply usage limits rather than a per-token bill. Check your service's terms.
Because it does not see letters, but pieces of words. Counting letters takes an extra reasoning step, and it sometimes slips. Always check it yourself.
| Term | What it means |
|---|---|
| Context window | The maximum number of tokens the model can consider at once: question, conversation and answer. |
| Model | The trained system that calculates, at each step, which token is most likely. |
| Probability | The score the model gives to each possible next token. |
| Temperature | A setting that controls how willing the model is to pick less likely tokens. |
| Token | A small piece of text (word, part of a word, punctuation) that AI reads and writes. |
| Tokenization | The step that cuts text into tokens and turns them into numbers. |
| Vocabulary | The closed list of tokens a model knows, each with its own numeric code. |
The token splits and probabilities in the examples are illustrative, chosen to explain the mechanism; real numbers depend on the model. The rules of thumb are orders of magnitude. Service prices and limits change often: always check the provider's current documentation. The course is adapted from the Italian original: the Italian examples were replaced with English ones.
Download the handoutPrintable PDF version of this lesson, with exercise and glossary.
Download the handout as PDF