The machine underneath: what a GPU is actually doing while your model runs
Almost everything else in this guide sits on top of this layer and never looks down at it. That is usually fine, right up until someone asks why inference gets cheaper per request when you batch it, why a model that “fits in 40 GB” does not fit, or why the expensive card is idle 70% of the time. Those are all the same question, and it has one answer: a GPU is enormously fast at arithmetic and comparatively slow at fetching the numbers to do arithmetic on. Nine sections, plain English first, with the arithmetic worked on a named machine so you can check it.
The reference machine. Where a number is needed, this lab uses an NVIDIA A100 80 GB (SXM): 108 streaming multiprocessors, 2.0 TB/s of memory bandwidth, 19.5 TFLOP/s of ordinary FP32 arithmetic and 312 TFLOP/s of FP16 tensor-core arithmetic. Different cards move the numbers; none of them change the shape of the argument. Where a figure is approximate or varies by generation, it says so.