NVFP4 vs FP8 vs MXFP4 vs EXL3: The Quantization Format Guide for Local LLMs
The same 320B model is 328 GB in FP8 and near 135 GB at 4 bits. Bytes-per-parameter math, per-format tradeoffs, and how to size a box that actually fits.
Tag archive
Coverage tagged nvfp4 quantization.
The same 320B model is 328 GB in FP8 and near 135 GB at 4 bits. Bytes-per-parameter math, per-format tradeoffs, and how to size a box that actually fits.