A use-after-free in llama-server lets a remote client crash it with one POST /completion request. Builds b8227 to b11392 are affected; upgrade to b11393.
A use-after-free in llama.cpp's HTTP server lets a remote client crash llama-server with a single POST /completion request. The CNA, VulnCheck, says the same flaw can be used to "shape a heap write primitive." Builds from b8227 up to, but not including, b11393 are affected. The fix shipped in b11393 on 4 October 2026. The CVE was published on 7 October.
The practical risk depends on exposure. llama-server listens on 127.0.0.1 by default, so a stock local install is not reachable from the network. Anyone who started it with --host 0.0.0.0, as the project's own Docker examples do, and did not set an API key should treat this as urgent. (We made the same point about the unauthenticated TensorFold server.)
What is confirmed
- Identifier: CVE-2026-107183, CWE-416 (use after free). VulnCheck is the CNA. NVD lists it as "Awaiting Analysis", and the CVSS scores below are VulnCheck's, not NVD's own.
- Scores: CVSS v4.0 9.2 (Critical),
CVSS:4.0/AV:N/AC:H/AT:N/PR:N/UI:N/VC:H/VI:H/VA:H/SC:N/SI:N/SA:N. CVSS v3.1 8.1 (High),CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H. Attack complexity is rated High in both. - Affected: llama.cpp builds >= b8227 and < b11393.
- Fixed: b11393, via pull request #29942 (commit
dbe4c3ed42343f0a7ba0fd7e808ffeaca404d29f), merged 4 October 2026 at 15:59 UTC. The release was published 23 minutes later. - Credit: the advisory credits "Anas", who also authored the fix.
How the bug works
The flaw is in common_chat_peg_mapper::map in common/chat-peg-parser.cpp, the code that turns parsed model output into tool calls. When the parser sees a TOOL_CLOSE event, it resets an optional that holds the pending tool call, which destroys the object inside. A raw pointer, current_tool, still points at that object. If a TOOL_ID event follows, the mapper writes through the dangling pointer into freed memory. The result is heap corruption, and in the reporter's test a double free.
The fix is one line: clear current_tool whenever pending_tool_call is reset.
Why it is remotely reachable
None of llama.cpp's built-in parsers emit a tool ID after a tool close. The maintainer who reviewed the pull request first called the change defensive for that reason. The reporter then pointed out that /completion accepts a client-supplied chat_parser (in tools/server/server-schema.cpp). A request carrying a parser with that ordering reaches the bad path. The reporter says a test request made the worker abort with free(): double free detected in tcache 2. The maintainer accepted the point. The reproducing request is posted in the public pull request comments, so the trigger is not secret. We do not reproduce it here.
NVD and VulnCheck describe the attacker as unauthenticated and the vector as network. By default llama-server has no API key, so on an exposed instance that matches the default setup. With a key set, the server's middleware skips only the health endpoints and UI assets (plus /models in older builds), so POST /completion requires the key in every affected build.
Remote code execution: claimed, not shown
The advisory and NVD stop at heap corruption and a write primitive. They do not claim code execution. In a comment on the pull request, the reporter says they chained this bug with a separate out-of-bounds read in the embeddings path to leak libc, and got remote code execution on a stock model under ASLR. They offered a write-up and proof of concept to the maintainers. We found no published write-up, and the claim is unverified. The embeddings read has no CVE in the sources we reviewed. Read the claim as: a crash is confirmed, a write primitive is asserted by the CNA, and full RCE needs a second bug and rests on one researcher's statement.
Exploitation status
As of 9 October 2026 the CVE is not in CISA's Known Exploited Vulnerabilities catalog. CISA's SSVC entry, attached to the NVD record on 7 October, reads exploitation "none," automatable "no," technical impact "total." We found no report of in-the-wild exploitation. The public request body in the pull request lowers the effort for a crash-only attack, so that status may change.
What defenders should do
- Check your build. Run
llama-server --version, or queryGET /propsand readbuild_info, which has the formb<build>-<hash>. Anything from b8227 to b11392 is affected. - Upgrade to b11393 or later. This is the only fix named by the advisory. Rebuild Docker images that pin an older tag.
- Find exposed servers. The default bind is 127.0.0.1. Check for
--host 0.0.0.0,LLAMA_ARG_HOST, published container ports, and reverse proxies that forward to the server. - Set a key where you cannot patch yet.
--api-key,--api-key-fileorLLAMA_API_KEYenable key authentication. With a key set,POST /completionrequires it in every affected build; also restrict the port to trusted networks. - Watch for crashes. Unexpected
llama-serveraborts or restarts, with heap errors such as "double free detected," on an internet-facing host are worth investigating.
People running llama.cpp locally for chat or coding assistants, for example on a Mac, are mostly unaffected while the server stays on loopback. Our comparison of local runtimes, MLX vs llama.cpp on Apple Silicon, covers when people choose llama.cpp's server.