tokens · attention · cache · precision

One token at a time: what actually happens between your prompt and the answer

A language model does exactly one thing: given some text, it produces a probability for every possible next token. Everything a chat assistant appears to do (reasoning, refusing, remembering what you said four messages ago) is that single operation run in a loop with its own output fed back in. This lab opens the loop up. Nine sections, each with something you can operate, and a running answer to the question that decides every production trade-off you will make: where does the memory go?

Two neighbours. The GPU lab explains why moving bytes, not doing arithmetic, is what costs you. This lab is where you find out which bytes. Neither is a prerequisite for the other, and the arithmetic here uses named models so you can check it.