How AI actually thinks
What you'll be able to do
- Explain that an LLM predicts the next token from probabilities instead of looking facts up
- Define token, context window, and variability in plain language with no math
- Predict why the same prompt wanders and why a long chat starts to forget
- Decide what front-loading context and trusting fluency mean for your Monday work
It is autocomplete with a college degree
An AI model predicts the most likely next chunk of text given everything that has come before. That single fact explains most of what feels like magic and most of what goes wrong. It is not reasoning the way you do, and it is not pulling answers out of a filing cabinet. It is guessing the next piece, very well, over and over.
Tokens, not words
The model does not read in whole words. It reads in tokens, which are word-pieces. Tokens are how length gets measured and how you get billed. You do not need the math, but you do need the instinct: a giant block of pasted text is not free, and very long inputs cost more and leave less room for the answer.
The context window is its short-term memory
Everything you paste plus everything the model has already said lives in one finite window. When that window fills up, the oldest material falls out of view. That is why a long chat seems to forget an instruction you gave at the top. Nothing broke. The early text simply scrolled past the edge of what the model can still see.
Probabilities, not facts
The model picks likely words, not true ones. A confident, well-formed sentence and a flatly wrong one are produced by the same machinery. There is no internal alarm that goes off when it is making something up. That is your job, and it is the whole reason you stay the editor.
Why the same prompt wanders
A bit of deliberate randomness keeps the model from sounding like a robot reading the same script. The side effect is that the same question can give you a slightly different answer each time. That is normal. It is not a sign the model changed its mind or learned something new.
What this means Monday
Three habits follow directly. Front-load your context so the model has what it needs early. Keep your most important instructions near the bottom of a long chat where they are freshest. And never treat a fluent paragraph as proof that the content is correct. The intern is fast and articulate. You still read every line.