AgentXploit exploits 72 previously disclosed agent-framework flaws at 59.3% end to end, against 38.4% for Codex on the same model.
A nine-author team (Weida Liang, Shi Qiu, Zhun Wang, Simon Sure, Xiaoyuan Liu, Tianneng Shi, Zhaorun Chen, Wenbo Guo and Dawn Song; the paper lists the National University of Singapore, UNC Chapel Hill, UC Berkeley, the University of Chicago and UC Santa Barbara) has published AgentXploit, an autonomous red-teaming system that reads an AI agent project's source code, proposes attack paths, and then tries them against a running copy. Posted to arXiv on 25 September 2026 (2609.31318), the paper reports 59.3% end-to-end success on a 72-instance benchmark, against 38.4% for a Codex CLI baseline using the same gpt-5.1-codex model. Anyone maintaining or deploying open-source agent frameworks is the audience: the same technique that audits them can be pointed at them.
One point matters before the numbers. The 72 "reproducible vulnerabilities" are not new discoveries. The authors built AgentXploit-Bench from publicly disclosed CVEs and agent-security issues from 2023 to 2025, including NVD records, and checked each against the vulnerable code and a runnable environment. The test is whether a system can rediscover and exploit a known flaw when the vulnerability details are withheld. The paper does not report new CVEs, and it does not describe a vendor disclosure process for new findings.
What the system does
AgentXploit has two roles. The Analyzer Agent maps the repository, keeps an investigation queue, and uses search, dependency, symbol, AST and LSP tooling to trace paths from external inputs to sensitive operations. Its output is a set of candidate attack paths with code evidence, treated as hypotheses. The Exploiter Agent takes those paths and attacks an isolated target through the attacker interface only. After each attempt it sees its previous payload, the target's response and the permitted execution trace, and revises. For prompt-injection tasks it starts from a curated seed corpus and can spread coordinated instructions across several injection locations. An external deterministic verifier, not the agent, decides whether the harmful outcome occurred.
The benchmark
The 72 instances span 12 open-source projects: GPT_Academic (22), OpenClaw (10), AgentScope (9), LangChain (8), LobeChat (8), LlamaIndex (4), AutoGPT (3), DB-GPT (2), GPT-Researcher (2), MetaGPT (2), OpenHands (1) and RAGFlow (1). The paper lists conventional weakness classes, namely path traversal, command execution, SQL injection, SSRF and unsafe evaluation, alongside LLM-mediated paths where adversarial content steers a tool call. Thirteen instances are indirect paths. One appendix example is CVE-2025-10236, a LaTeX path traversal.
Results
- End-to-end, averaged over three runs: 59.3% (plus or minus 2.1) for AgentXploit versus 38.4% (plus or minus 2.9) for Codex, or 128 versus 83 successes out of 216 runs. The gap is 30.8 points on indirect paths and 18.6 on direct paths.
- On AgentDojo, where injection points are supplied, the Exploiter Agent reaches 79.2%, versus 52.7% for AgentVigil and 41.5% for handcrafted attacks.
- Compute explains part of the gap. AgentXploit averages about 108k tokens and 455 seconds per task; default Codex averages 29.5k tokens and 62 seconds. Giving Codex 4x tokens lifts it to 45.8%, and matching AgentXploit's budget gives 46.3%, leaving a 13.0-point difference.
Caveats from the paper
- GPT_Academic supplies 30.6% of the instances. Without it, AgentXploit succeeded on 24 of 50 instances and Codex on 14 of 50 in the initial run.
- The set is curated for reproducibility and is not a random sample of repositories.
- Precision is modest. On OpenHands, two of five candidates were valid in-scope vulnerabilities, one was a real out-of-scope finding, and two were false positives.
- The authors do not claim the evaluation is free of contamination, since the backbone model may have seen the public CVEs. A code-mutation test that renames identifiers and strips comments is meant to narrow that concern.
What defenders should do
- Treat the benchmark's classes as the baseline for your own agent stack. Path traversal, command execution, SSRF and unsafe evaluation reachable from model-controlled tool arguments are the recurring shapes.
- Assume untrusted retrieved content can reach file, shell and network tools, and enforce destination and path allow-lists in the tool layer rather than in prompts.
- Patch the 2023 to 2025 CVEs in the frameworks you run. These are the flaws the system was built to rediscover.
- Run source-aware automated auditing against your own agent deployments, and expect lower-effort attackers to have access to the same approach. The authors released evaluation artifacts and withheld reusable payload text from the manuscript under a responsible-release process.