vLLM v0.30.0 Upgrade Notes: Fast Start, HiSparse, New Models and the Breaking Changes
vLLM v0.30.0 adds DeepSeek-V4.1-Flash, GLM-5.3-Flash, Fast Start and HiSparse, and changes defaults that can break an upgrade. What to check.
· 4 minDesk · Labs & reverse engineering
Technical deep dives, reproducible tests, and tool evaluations.
vLLM v0.30.0 adds DeepSeek-V4.1-Flash, GLM-5.3-Flash, Fast Start and HiSparse, and changes defaults that can break an upgrade. What to check.
· 4 minMiniMax's M3.1 Flash preview brings a 1M-token context, image and video input and five effort levels to MiniMax Code. What it does, and how it tests.
· 4 minChina Telecom's Apache 2.0 MoE, trained on Huawei Ascend, scores 75.00 on SWE-bench Verified in its own table, a point behind Qwen3.6, and its 4-bit file is ~18 GB.
· 6 minOn 19 Oct GitHub Copilot drops GPT-5.5, GPT-5.4, two mini models, Gemini 3.7 Flash and Grok 4.5. Here is each replacement, what it costs and what admins must do.
· 7 minSonnet 4.5 stops answering on Nov 30, 2026. Sonnet 5.5 is a third cheaper per token, but prefill, forced tool use and thinking budgets now return 400 errors.
· 8 minThe leaked 'Gemini 4 Pro' was billed as an Astra and Opus killer. The real Gemini 4 Argon wins Google's tests, ties Astra on neutral ones, and is cheap for now.
· 6 minAnt Group's 560B Ling 3.1 Flash lands against GLM 5.3 Flash and Qwen 3.8 Flash Next. Specs, licences, prices, and the benchmarks Ling has not published yet.
· 6 minAnt's 560B-parameter Ling-3.1-flash beats rival flash models on SWE-Pro and posts 87.90 on CyberGym, but it ships with no safety evals and no weights yet.
· 6 min