The engine, not the model: why serving is a systems problem wearing a machine-learning hat
You can run a model in about four lines of Python. Nobody serves one that way, and the gap between those four lines and a production endpoint is almost entirely systems work: deciding what to run each step, what to keep in memory, what to throw away, what to gamble on, and which of your users to make wait. This lab is that machinery, nine sections, each with something you can operate, and a running argument that the number your users feel is not the number your dashboard shows.
Where this sits. The GPU lab explains why moving bytes decides everything; the LLM lab explains which bytes. This one is about the machinery built to manage them. None is a prerequisite for the others, and the mechanisms here are described generally, vLLM, SGLang and TensorRT-LLM implement them differently, and are named where the difference matters.