A model that fits on paper can still fail when real users arrive. The missing line is often not the model weights: it is the key-value cache, runtime workspace, or a parallel layout that cannot distribute memory as evenly as the spreadsheet assumes.
This GPU memory sizing guide answers a practical procurement question: how many H100 or H200 GPUs should you budget for a 70-billion-parameter inference service? It includes a reusable calculation, a concurrency table, two original diagrams, and an acceptance checklist. All designs are hypothetical; the arithmetic was executed in Python, but no GPU benchmarks or deployment tests were run for this article.