NVFP4 vs FP8 vs MXFP4 vs EXL3: The Quantization Format Guide for Local LLMs
The same 320B model is 328 GB in FP8 and near 135 GB at 4 bits. Bytes-per-parameter math, per-format tradeoffs, and how to size a box that actually fits.
· 13 minDesk · Labs & reverse engineering
Technical deep dives, reproducible tests, and tool evaluations.
The same 320B model is 328 GB in FP8 and near 135 GB at 4 bits. Bytes-per-parameter math, per-format tradeoffs, and how to size a box that actually fits.
· 13 minPrepare for the September 2026 FIPS 140-2 sunset. Learn the procurement impacts, technical breaking changes, and exact steps to migrate stacks to FIPS 140-3.
· 8 minExplore the EU Cyber Resilience Act Article 14 mandatory 24-hour vulnerability reporting rules, ENISA platform architecture, and PSIRT compliance steps.
· 9 minHow to detect and evict CosmicSting (CVE-2024-34102) backdoors from Magento and Adobe Commerce, plus a correction: StyleSmuggler is a separate zero-day, CVE-2026-75650.
· 9 minCritical GitLab arbitrary file read flaw leaks server secrets. Learn how to hunt logs, stop admin session forgery, and execute an incident recovery runbook.
· 9 minInterlock ransomware targets Cisco FMC flaw CVE-2026-20131 to breach networks. Discover why standard patching fails and how to audit and remove backdoors.
· 7 minA stability-first SGLang deployment serves DeepSeek V4.1 Flash across three NVIDIA DGX Sparks: 750k KV cache, ~1600 tok/s prefill, ~38 tok/s single-stream decode.
· 2 minDeploying DeepSeek-R1 agent runtimes? Learn how to harden the harness architecture, configure local inference, mitigate CVEs, and enforce sandbox security.
· 7 min