precision · recall · ranking · drift · the judge

The number to trust: evaluation metrics when a miss costs more than a false alarm

One story runs through this whole lab. A chain of learning centres uses an AI screener that reads teachers’ observation notes and drafts reports for parents, flagging which learners need reading support. Miss a learner and they lose a term of help; flag the wrong one and a teacher loses twenty minutes. Every metric decision on this page falls out of that one asymmetry, and every simulation uses the same class of twenty so the numbers line up across tabs. Nothing here needs machine-learning background: if you can read a school report, you can run this lab.

One neighbour. The evaluation lab teaches the process: rubrics, judges, golden sets, the improvement loop. This lab teaches the numbers themselves, which one answers which question, and the traps between them. Either order works.