loss · gradient descent · overfitting · knowing when to stop

Wrong, then less wrong: watch a model learn

Every other lab in this track starts after training is finished: the GPU lab runs a trained model, the LLM lab opens one up, serving and fleet put one behind an endpoint. This lab is the missing step before all of them. A model starts wrong, measures exactly how wrong, and nudges every number it owns in the direction that makes it slightly less wrong, millions of times. Nothing on this page is a recording: the networks here train live, in this tab, while you watch.

Two neighbours. The GPU lab explains why each nudge is cheap only when you do thousands at once, and the LLM lab is what one of these looks like grown to billions of numbers. This page is the same mechanism small enough to see.