high Cve 2026 107183 · AI Security

llama.cpp Server Bug CVE-2026-107183: A Crafted Chat Parser in One Request Corrupts the Heap

Data graphic: Poster for llama.cpp server bug CVE-2026-107183, a use after free CWE-416 in llama-server, rated CVSS v4.0 9.2 critical. Builds b8227 to b11392 are affected, the fix is b11393, and no exploitation is known.
SH

Vulnerability analyst · Published Oct 8, 2026, 9:41 PM EDT

A use-after-free in llama-server lets a remote client crash it with one POST /completion request. Builds b8227 to b11392 are affected; upgrade to b11393.

A use-after-free in llama.cpp's HTTP server lets a remote client crash llama-server with a single POST /completion request. The CNA, VulnCheck, says the same flaw can be used to "shape a heap write primitive." Builds from b8227 up to, but not including, b11393 are affected. The fix shipped in b11393 on 4 October 2026. The CVE was published on 7 October.

The practical risk depends on exposure. llama-server listens on 127.0.0.1 by default, so a stock local install is not reachable from the network. Anyone who started it with --host 0.0.0.0, as the project's own Docker examples do, and did not set an API key should treat this as urgent. (We made the same point about the unauthenticated TensorFold server.)

What is confirmed

  • Identifier: CVE-2026-107183, CWE-416 (use after free). VulnCheck is the CNA. NVD lists it as "Awaiting Analysis", and the CVSS scores below are VulnCheck's, not NVD's own.
  • Scores: CVSS v4.0 9.2 (Critical), CVSS:4.0/AV:N/AC:H/AT:N/PR:N/UI:N/VC:H/VI:H/VA:H/SC:N/SI:N/SA:N. CVSS v3.1 8.1 (High), CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H. Attack complexity is rated High in both.
  • Affected: llama.cpp builds >= b8227 and < b11393.
  • Fixed: b11393, via pull request #29942 (commit dbe4c3ed42343f0a7ba0fd7e808ffeaca404d29f), merged 4 October 2026 at 15:59 UTC. The release was published 23 minutes later.
  • Credit: the advisory credits "Anas", who also authored the fix.

How the bug works

The flaw is in common_chat_peg_mapper::map in common/chat-peg-parser.cpp, the code that turns parsed model output into tool calls. When the parser sees a TOOL_CLOSE event, it resets an optional that holds the pending tool call, which destroys the object inside. A raw pointer, current_tool, still points at that object. If a TOOL_ID event follows, the mapper writes through the dangling pointer into freed memory. The result is heap corruption, and in the reporter's test a double free.

The fix is one line: clear current_tool whenever pending_tool_call is reset.

Why it is remotely reachable

None of llama.cpp's built-in parsers emit a tool ID after a tool close. The maintainer who reviewed the pull request first called the change defensive for that reason. The reporter then pointed out that /completion accepts a client-supplied chat_parser (in tools/server/server-schema.cpp). A request carrying a parser with that ordering reaches the bad path. The reporter says a test request made the worker abort with free(): double free detected in tcache 2. The maintainer accepted the point. The reproducing request is posted in the public pull request comments, so the trigger is not secret. We do not reproduce it here.

NVD and VulnCheck describe the attacker as unauthenticated and the vector as network. By default llama-server has no API key, so on an exposed instance that matches the default setup. With a key set, the server's middleware skips only the health endpoints and UI assets (plus /models in older builds), so POST /completion requires the key in every affected build.

Remote code execution: claimed, not shown

The advisory and NVD stop at heap corruption and a write primitive. They do not claim code execution. In a comment on the pull request, the reporter says they chained this bug with a separate out-of-bounds read in the embeddings path to leak libc, and got remote code execution on a stock model under ASLR. They offered a write-up and proof of concept to the maintainers. We found no published write-up, and the claim is unverified. The embeddings read has no CVE in the sources we reviewed. Read the claim as: a crash is confirmed, a write primitive is asserted by the CNA, and full RCE needs a second bug and rests on one researcher's statement.

Exploitation status

As of 9 October 2026 the CVE is not in CISA's Known Exploited Vulnerabilities catalog. CISA's SSVC entry, attached to the NVD record on 7 October, reads exploitation "none," automatable "no," technical impact "total." We found no report of in-the-wild exploitation. The public request body in the pull request lowers the effort for a crash-only attack, so that status may change.

What defenders should do

  1. Check your build. Run llama-server --version, or query GET /props and read build_info, which has the form b<build>-<hash>. Anything from b8227 to b11392 is affected.
  2. Upgrade to b11393 or later. This is the only fix named by the advisory. Rebuild Docker images that pin an older tag.
  3. Find exposed servers. The default bind is 127.0.0.1. Check for --host 0.0.0.0, LLAMA_ARG_HOST, published container ports, and reverse proxies that forward to the server.
  4. Set a key where you cannot patch yet. --api-key, --api-key-file or LLAMA_API_KEY enable key authentication. With a key set, POST /completion requires it in every affected build; also restrict the port to trusted networks.
  5. Watch for crashes. Unexpected llama-server aborts or restarts, with heap errors such as "double free detected," on an internet-facing host are worth investigating.

People running llama.cpp locally for chat or coding assistants, for example on a Mac, are mostly unaffected while the server stays on loopback. Our comparison of local runtimes, MLX vs llama.cpp on Apple Silicon, covers when people choose llama.cpp's server.

Sources

Keep reading

All latest →
  1. highAI SecurityFake ChatGPT, Gemini, Claude and Muse ad portals use a fake browser window to steal ad-account logins7 min
  2. watchAI SecurityOpenAI Notified 100+ Organizations About Model Activity: A Notice Is Not a Compromise8 min
  3. highAI SecurityMCP TypeScript SDK Lets a Malicious Server Pull OAuth Secrets From Clients (CVE-2026-104850)4 min
  4. elevatedAI SecurityMistral Large 4 preview ships a reduced-moderation cyber tier, and Artificial Analysis lists its 82% as the top score6 min
  5. watchAI SecurityNextChat proxy fallback lets unauthenticated callers make the server fetch any URL (CVE-2026-105238)7 min
  6. highAI SecurityDify CVE-2026-105762: Unauthenticated SSRF in Remote-File Upload Reaches Internal Services and Cloud Metadata5 min