Skip to content
  • fundamentals

How Language Models Actually Work

Four mechanical facts about what happens between your prompt and the answer. Each one explains a behavior you have already noticed and probably found strange.

5 min read

What AI Actually Is covers what the thing is. This one covers what it does, mechanically, between you pressing enter and the words appearing.

It is not a systems diagram. There are four facts here, and each one exists in this post because it explains something you have already run into and filed under "the AI is being weird". Once you know the mechanism, the weirdness stops being random and starts being predictable, which is the whole point.

1. It does not read words

Before anything else happens, your text is chopped into tokens. A token is usually a chunk of a word rather than a word. For ordinary English, a token averages about four characters, so a thousand words comes out around thirteen hundred tokens.

The model never sees your sentence. It sees a list of these chunks.

This one fact explains an entire family of failures that otherwise look absurd:

  • Counting letters. Ask how many times "r" appears in "strawberry" and a model can get it wrong, because it is not looking at ten letters. It is looking at two or three chunks that it has learned usually mean strawberry. The letters are not individually there to count.
  • Spelling games, anagrams, "words ending in -ight". Same reason. You are asking about a level of detail below the one it operates at.
  • Character counts. "Write exactly 100 characters" is a request in a unit the model cannot perceive directly.

2. It writes one token at a time, with no plan

The model does not compose an answer and then type it out. It predicts the single most plausible next token, appends it, then looks at everything including that new token and predicts the next one. Over and over, until it predicts a token meaning "done".

There is no draft. There is no outline it is working from. Each token is chosen with only the tokens before it to go on.

It is not writing an answer. It is repeatedly extending a piece of text that increasingly looks like an answer.

That sounds like a technicality. It is the most practically useful fact in this post, because three things follow directly from it:

Its first few tokens commit it. If a reply opens with "Yes, there are three main approaches", it now has to produce three approaches, because that is what the text it has already written implies comes next. If there are really only two, you get a padded third. The answer was shaped by its own opening before it had considered your question fully.

Asking for reasoning genuinely helps. "Think it through step by step" is not flattery or a magic phrase. Reasoning tokens are more text in the window, and every subsequent token gets to condition on them. You are giving it room to do work before it commits to a conclusion, rather than making it commit in the first five words.

Its confidence is unrelated to its accuracy. A fluent, well-structured paragraph and a fabricated one are produced by the same process at the same level of certainty. That is the mechanism behind why it makes things up.

3. It rolls dice

At each step the model does not have one candidate token. It has a probability distribution over all of them, and it picks from that distribution. Most tools expose a temperature setting that controls how much it favors the likeliest option versus occasionally taking a less likely one.

This is why the same prompt gives you a different answer the second time. It is not the model changing its mind or having an off day. It is sampling.

Two practical consequences:

  • Getting a bad answer does not mean the prompt was bad. Try it again before you rewrite it. If you get a good answer one time in three, the prompt is probably fine and you were unlucky.
  • A good answer is not repeatable evidence. If you are choosing between two prompts, compare a few runs of each. One sample of each tells you almost nothing.

4. It stopped learning before you met it

The weights that make the model what it is were fixed at the end of training. Nothing you type changes them. This has three effects worth holding onto:

It has a cutoff. There is a date after which it has no training knowledge. It may still answer confidently about later events, because producing a plausible-looking answer is what it does. Some tools bolt on live search to cover this, which helps and is also a different mechanism with its own failure modes.

It does not remember you between conversations unless the product you are using has built a memory feature that quietly re-inserts things into your context. When it does remember, that is the product doing it, not the model.

Correcting it does not teach it anything. Explaining a mistake fixes the current conversation, because your correction is now in the context and the next tokens condition on it. Open a new chat and the same mistake is available again. If a correction matters, it belongs in your prompt or in a saved instruction, not in a conversation you will close.

What actually changes

Put together, these four facts move you from arguing with the model to operating it.

| What you noticed | What is happening | | --- | --- | | It miscounts letters | It never saw letters, only chunks | | It padded an answer to three points | Its own opening sentence committed it | | Same question, different answer | It samples rather than picks | | It repeated a mistake in a new chat | Nothing you said changed the model | | It was confidently wrong | Fluency and accuracy come from one process |

None of this makes the tool worse. It makes it legible. A model that is extending text one chunk at a time from frozen weights is very good at some things and structurally incapable of others, and knowing which is which is most of the skill.

The short version

It reads chunks, not words. It writes one chunk at a time with no plan, so its first sentence constrains the rest. It samples, so it is not deterministic. It learned nothing after training, so corrections do not persist.

Everything else you have read about prompting is downstream of those four. Context Is Everything covers the one you have the most control over.

Newsletter

The email list opens soon

I'm still setting up the newsletter. Subscribe on YouTube in the meantime and I'll announce it there first.

Related reading

All tutorials