discovery · routing · state transfer · control loops

Twenty engines, one endpoint: the layer that decides which of them gets your request

The serving lab is about one engine replica deciding what to run this step. Everything changes at twenty replicas. Which worker gets this request? Does any of them already hold the first two thousand tokens of it? Is it cheaper to move that state across the network or recompute it? Do prefill and decode scale together, or separately? And when a worker dies mid-stream, what does the person watching the text actually see? None of that happens inside an engine, and all of it decides whether the fleet is affordable. Eight sections, each with something you can operate.

Where this sits. Engines (vLLM, SGLang, TensorRT-LLM) schedule work inside one replica. This lab is the layer around them, the part usually called an inference gateway, router or orchestrator. The mechanisms are described generally; named systems appear where a concrete implementation makes the idea clearer.