Inference Engineering: A Practical Primer for Serving Your Own Models
Prefill is compute-bound, decode is bandwidth-bound, and KV cache is what really caps concurrency. The metrics, memory maths and tradeoffs behind serving.
Tag archive
Coverage tagged inference engineering.
Prefill is compute-bound, decode is bandwidth-bound, and KV cache is what really caps concurrency. The metrics, memory maths and tradeoffs behind serving.