Under the hood
Token by Token
How a chatbot writes its reply
Ask a real AI model a question. Then step through how its reply was made, one token at a time, using the real data behind it, with every value labeled.
The options for the next token
What you’ll see
Tokens
Text is split into Token: A piece of text the model works with: often a whole common word with its leading space, sometimes part of a word, a single character, or a byte. Like building blocks: a common word is usually one block, while a rare word is built from several.: often a whole word with its leading space, sometimes a piece of one, like building blocks of text. AI services count tokens, not words: a model can take in Context limit: The most tokens a model accepts in one request. This app keeps well under it with a smaller budget of its own. Like a page limit on what can be handed over at once., each reply is capped at a set number, and every request is priced by the token.
See it in the sample- "Why"13903
- space, "is"382
- space, "the"290
- space, "sky"17307
- space, "blue"9861
- "?"30
Calculated tokens and IDs, split by this app’s tokenizer
A weighted random pick
At every step the model Score: What the network produces for every possible next token. The scores are turned into percentages. Like a rating for every word in the dictionary, at every step: the higher the rating, the likelier that word comes next. the options for the next token. One is Weighted random pick: How one token is chosen from the options: at random, with likelier options picked more often. Also called sampling. Like a raffle where each option holds tickets in proportion to its chance: the favorite usually wins, but not always., like a raffle where likelier options hold more tickets, so the top option doesn’t always win.
See it in the sample- space, "Dust"
- space, "and"
- space, "tiny"
- space, "particles"
- space, "can"
- ␣intens39.4%
- ␣scatter26.2%
- ␣enhance16.2%picked
- ␣deepen5.57%
Recorded tokens, reported by OpenAICalculated percentages, from its logprobs
Context, not memory
With each message, this app sends the Context: Everything sent with a request: instructions, the conversation so far, and the new message. The model has no other memory of the chat. Like handing an actor the whole script so far before every new line: all the model has to go on is what’s in the script. again, with its own Instructions (system prompt): Text an app sends to the model ahead of your message, setting how it should behave. Many apps keep theirs hidden; this app shows its own. Like a briefing handed over before the conversation starts., like handing over the whole transcript each time. Chatting doesn’t change the model.
See it in the sample- InstructionsYou are the assistant on Token by Token, an educational website that shows how language models generate text one token at a time. Answer helpfully, accurately, and concisely, usually in under 150 words and in plain language. Use Markdown only when it helps. If you aren't sure about something, say so. You have no tools, no internet access, and no memory of other conversations.80 tokens
- YouWhy is the sky blue?6 tokens
- AssistantSunlight contains many colors, but air molecules scatter shorter wavelengths more than longer ones. Blue light has a shorter wavelength than red light, so it gets scattered in all directions across the sky and reaches your eyes from everywhere. That’s why the daytime sky looks blue.53 tokens
- YouWhy are sunsets red, then?7 tokens
Recorded text, sent by this appCalculated token counts
What’s real here
Every value carries one of four labels, always as a word, never just a color:
- Recorded
- Sent, done, or measured by this app, or reported by OpenAI, for your conversation.
- Calculated
- Computed here from recorded values, with the method named.
- Reference
- Documented facts, such as the model’s context limit, with a source and a date.
- Example
- A teaching drawing, not measured from the model. Views of the network’s insides are always examples: OpenAI hasn’t published the design of its hosted models.
A What-if tag marks anything simulated on the page. The model wasn’t asked again.