All guides

Fundamentals · Intermediate

How AI Assistants Actually Generate Answers

Tokens, context windows, temperature, and why an AI assistant sounds equally confident whether it is right or wrong. The mechanism, explained without maths.

Simanta Pratim DasPublished 7 min read

You do not need to understand transformers to use an AI assistant well. But you do need a working model of the mechanism, because almost every frustrating behaviour — forgetting things, inventing citations, giving different answers to the same question — becomes predictable once you know how the output is produced.

This guide covers the four concepts that explain most behaviour: tokens, the context window, next-token prediction, and sampling.

Tokens: the model never sees words

Before anything happens, your text is split into tokens — chunks that are usually part of a word.

"Understanding tokenisation helps."
→  ["Under", "standing", " token", "isation", " helps", "."]

Rough rule: one token ≈ 4 characters of English, so 100 words is about 130 tokens. Code, unusual names, and non-English text fragment into more tokens.

This single detail explains a family of otherwise baffling failures:

TaskTypical resultWhy
"How many r's in strawberry?"Often wrongIt sees straw + berry, not letters
"Reverse this string"UnreliableCharacter order is not visible at token level
"Write a 500-word essay"ApproximateIt cannot count words while generating
Rhyming in some languagesWeakRhyme lives in sounds, not token boundaries

So when an assistant fails at counting letters, it is not "dumb" — it is being asked about a level of detail that was discarded before the model saw your message. Ask it to write a script to count the letters instead and it will get it right, because now the task is code generation rather than character inspection.

The context window: there is no memory

This is the most important and least intuitive part.

The model has no memory between messages. None. What creates the illusion of a conversation is that the entire visible history is resent as one block of text on every single turn.

Turn 5 of a conversation is not "the model remembering turns 1–4." It is a fresh request containing turns 1–4 pasted in as context, plus your new message. The model reads all of it from scratch, every time.

The context window is the ceiling on how much text fits in that block — typically tens to hundreds of thousands of tokens depending on the model. Everything must fit inside it: system instructions, your conversation, any attached documents, and the answer being generated.

When a conversation exceeds the window, something has to go. Usually the oldest messages are dropped or compressed. So:

  • "It forgot what I said earlier" — that text is no longer in the input. It is not buried; it is absent.
  • "It contradicted itself" — the earlier statement fell out of the window.
  • Long conversations degrade — more context means more competing material, and relevant details get diluted. Models also attend less reliably to the middle of very long inputs than to the beginning and end.

The practical habit: start a fresh conversation when you switch topics, and restate critical constraints rather than assuming they persist. Long threads are not free. If a detail matters twenty messages later, say it again.

This is also why "memory" features in AI products are not the model remembering. They are a separate store of facts about you that gets re-injected into the context on each request. Useful, but architecturally a different thing.

Next-token prediction: one word at a time, no plan

The model does exactly one thing: given all the text so far, produce a probability distribution over what token comes next.

Given "The capital of France is", it might produce:

" Paris"      92%
" located"     3%
" the"         2%
" a"           1%
...

One token is selected. It is appended. The whole thing runs again with the new text included. Repeat until a stopping point.

Two consequences that matter enormously:

1. There is no plan. The model does not outline the answer and then write it. It commits to token 1 before knowing token 50. Coherence emerges because each token is chosen in light of everything before it — not because anything was decided in advance.

This is why "show your reasoning before concluding" genuinely improves accuracy. It is not a psychological trick. Reasoning tokens generated first become part of the input for the tokens that follow, so the conclusion is computed with the analysis available. A conclusion stated first cannot be informed by reasoning that comes after it.

2. There is no truth check. Nowhere in this loop is there a lookup against a database of facts. "Plausible" and "true" are different things that usually coincide, because training text is mostly accurate — but when they diverge, the model follows plausibility.

Where fabricated citations come from

This mechanism explains the classic failure precisely. Ask for a source and the model generates:

" Smith"  →  " ,"  →  " J"  →  "."  →  " ("  →  "201"  →  "9"  →  ")"  ...

Every token is a highly plausible continuation of an academic reference. The format is learned perfectly. But there is no step that checks whether Smith (2019) exists. A fake citation is not a malfunction — it is the system working exactly as designed on a task it was never designed for.

This is why the fix is never better prompting. It is retrieval: make the model read a real source rather than recall one.

Sampling and temperature: why answers vary

If the highest-probability token were always chosen, output would be repetitive and oddly flat. So selection involves controlled randomness, governed by temperature.

TemperatureBehaviourSuits
Low (0–0.3)Nearly deterministic, picks the likeliest tokenExtraction, code, classification, factual answers
Medium (0.5–0.8)Balanced — typical chat defaultGeneral writing, explanation
High (1.0+)Adventurous, less predictableBrainstorming, creative variation

This is why the same question can produce different answers, and why re-asking is not verification — you are drawing a second sample from the same distribution, not getting a second opinion. If a model is confidently wrong about something, it will often be confidently wrong again in slightly different words.

It also explains why "Are you sure?" is useless as a check. Agreement is a plausible continuation of being challenged. Models frequently abandon correct answers under mild pressure. You are testing agreeableness, not accuracy.

Why confidence is uninformative

Here is the uncomfortable part. The model's internal probabilities are about token likelihood, not factual certainty. And the confident register — declarative sentences, no hedging — is a writing style learned from training data, where reference material is written confidently.

So:

  • Fluency does not indicate accuracy.
  • Absence of hedging does not indicate certainty.
  • Presence of hedging does not reliably indicate genuine uncertainty either.

There is no honest confidence signal in the text. When an assistant says "I'm not certain, but..." that phrase was generated by the same plausibility mechanism as everything else. Treat calibration language as style, not as data.

What retrieval and search actually change

When an assistant "searches the web," the loop does not change. What changes is the input:

  1. Your question is turned into a search query.
  2. Real pages are fetched.
  3. Their text is inserted into the context window.
  4. The same next-token prediction runs — but now with real source text present.

This is a genuine reliability improvement, because plausible continuations of a real document tend to be accurate. It is the same reason pasting your own source material works so well.

But it is not a guarantee. The model can still misread a source, blend two sources, or summarise a low-quality page confidently. Retrieval changes where the facts come from; it does not add a verification step. Always check the cited source actually says what the summary claims.

How this maps onto PhantomAI

PhantomAI's routing is a direct application of the above. Two models are available: openai/gpt-oss-20b for quick exchanges where speed matters, and openai/gpt-oss-120b for reasoning-heavy work — both served through Groq. Requests pass through a small pipeline that classifies the task, calls the model, and formats the result.

When a question needs current information, PhantomAI searches and puts retrieved text into the context. When you attach a document or URL, the same thing happens with your material. That is the highest-reliability mode available and the one we would encourage you to default to for anything factual.

Everything on this page still applies. There is no configuration in which the mechanism stops being next-token prediction. See our limitations page for the specific consequences.

Five habits that follow from the mechanism

  1. Supply sources instead of relying on recall. Converts fabrication risk into reading comprehension.
  2. Ask for reasoning before conclusions. Mechanically improves multi-step accuracy and gives you something auditable.
  3. Start fresh threads for new topics. Context is finite and dilution is real.
  4. Never treat re-asking as verification. Check the source, not the model.
  5. Ignore confident tone entirely. It carries no information about accuracy.

Continue

Simanta Pratim Das

Founder & Developer

Simanta is an independent AI engineer based in Guwahati, India, building PhantomAI as a solo project — designing the product, the interface, and the AI pipeline end to end.

Related guides

Try these techniques in PhantomAI

PhantomAI is free to use, and the workflows in this guide work best when you paste in your own material rather than relying on the model's memory. Before you start, it's worth reading what it can't do.