> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

ExploitGym: Teaching AI Agents to Hack With Real-World CVEs

[ View on GitHub ]

ExploitGym: Teaching AI Agents to Hack With Real-World CVEs

Hook

An AI agent that can write a working remote code execution exploit from a CVE description would be worth billions—or terrifying, depending on who controls it. ExploitGym is the first benchmark designed to measure exactly that capability.

Context

Security researchers have been evaluating LLM capabilities on offensive tasks for years, but their benchmarks share a fatal flaw: they test vulnerability discovery, not exploitation. Finding a buffer overflow is freshman-level work; turning it into working shellcode that bypasses ASLR, NX, and stack canaries is where expertise lives. Existing benchmarks like CWE-based synthetic datasets or CTF archives don't cut it—synthetic challenges lack the dependency chaos and incomplete debug symbols of real-world targets, while CTF challenges are designed for human intuition with writeups that assume you already know to "check the heap metadata."

The ExploitGym team at UC Berkeley's SUNBLAZE lab built something different: a containerized evaluation harness containing 869 real CVE instances spanning userspace C programs, the V8 JavaScript engine, and Linux kernel vulnerabilities. Each task is frozen at the exact vulnerable commit with period-accurate dependencies, creating a time capsule of exploitation conditions. The architecture enforces realistic constraints—network isolation, metered LLM calls, no internet access—while preventing agents from cheating their way to flags. This isn't a CTF platform; it's infrastructure for answering whether autonomous agents can do the hardest part of offensive security.

Technical Insight

Isolation

1. Request Task

2. Generate Secrets

Token Salt, Flag Seed

3. Spawn Container

SSH/Network Endpoint

4. Exploit Binary

Derived Flag

5. Submit Flag

6. Validate via Seed

Enforce Network Rules

Allow LLM Access

Metered API Calls

Agent/LLM

Controller Service

Squid Firewall Proxy

LLM Proxy

Ephemeral Docker Container

Cryptographic Secrets

System architecture — auto-generated

The core architectural insight is separating task provisioning from secret generation. When the controller starts, it generates three cryptographic values—a token salt for API authentication, a flag seed for deterministic verification, and session keys—that never touch version control. Flags aren't stored in containers; they're derived server-side from the seed and task ID. An agent that compromises a container and reads every file will find nothing.

Here's the flag derivation mechanism from the controller code:

# Controller generates seed on startup
flag_seed = secrets.token_bytes(32)

# For each task instance
def generate_flag(task_id: str, seed: bytes) -> str:
    h = hashlib.sha256()
    h.update(seed)
    h.update(task_id.encode())
    return f"flag{{{h.hexdigest()[:32]}}}"

# Agent must derive flag from exploitation artifact
# e.g., reading /root/flag.txt after privilege escalation
# Container has placeholder that agent replaces via exploit

This design forces agents to actually exploit the vulnerability. A userspace heap overflow must achieve arbitrary write to replace the placeholder. A kernel race condition must escalate to root. A V8 type confusion must achieve code execution within the renderer process. Reading files isn't enough—you need primitive control.

The three-tier isolation model creates realistic offensive constraints. Layer one is Docker containers—each task runs in a purpose-built image with the vulnerable binary, exact library versions (glibc 2.23 for older challenges, modern versions for recent CVEs), and a frozen filesystem. Layer two is the Squid-based firewall proxy. Agents receive network access to their allocated container and the LLM proxy, nothing else. No GitHub for downloading exploit frameworks, no Shodan for reconnaissance, no VirusTotal for checking detection rates. Layer three is the LLM proxy itself, which meters API calls and logs prompts. This prevents agents from burning unlimited tokens on brute-force approaches or exfiltrating task details through prompt injection.

Task diversity spans three domains that require fundamentally different mental models. Userspace C challenges include heap corruption (UAF, double-free, heap overflow), stack smashing with canary bypasses, and format string vulnerabilities. These require understanding glibc malloc internals, ROP chain construction, and ASLR defeats through info leaks. V8 tasks involve JIT engine bugs—type confusions, bounds check eliminations, and prototype chain corruptions. Exploiting these demands knowledge of JavaScript engine internals, object layouts, and how to turn a type confusion into arbitrary read/write. Kernel challenges are race conditions, use-after-frees in device drivers, and privilege escalation bugs. Success requires understanding kernel memory management, synchronization primitives, and how to weaponize a race window.

The evaluation flow uses a REST API where agents request tasks, receive connection details, and submit flags. For SSH-based tasks, the controller provisions a container and returns credentials:

# Agent requests task
POST /api/task/request
{"task_id": "CVE-2016-4484_crypto_lab"}

# Controller response
{
  "container_id": "exploit-gym-a3f891",
  "ssh_host": "172.16.0.15",
  "ssh_port": 2222,
  "username": "ctf",
  "password": "temp_9x3k2",
  "timeout": 3600
}

# Agent exploits, derives flag, submits
POST /api/flag/submit
{"task_id": "CVE-2016-4484_crypto_lab", "flag": "flag{a3d8f...}"}

For network-service tasks (HTTP servers, daemon processes), agents receive an IP and port but no shell access. This mirrors real penetration testing—you have to achieve remote code execution from outside the system.

The use of uv for Python dependency management and bundled static binaries (GDB, socat, node) is subtle but critical. Most security benchmarks fail at reproducibility—one evaluator has GDB 10 with Python 3 support, another has GDB 7 without, and suddenly exploits that rely on scripted heap introspection break. ExploitGym bundles exact tool versions, ensuring that a successful exploit in one environment works in another. This also prevents agents from relying on external tooling that might not exist in real engagements.

Gotcha

The operational overhead is substantial. Running ExploitGym requires four separate services (controller, runner, firewall proxy, LLM proxy), a Docker daemon with significant disk space (images total ~50GB for the full task set), and coordination of shared secrets between components. There's no lightweight mode—you can't just run a single task in a subprocess for rapid iteration. If you're developing an agent, the debug cycle is: modify agent code, restart controller to pick up new secrets, provision container, wait for Docker pull, attempt exploit, collect logs from three different services. This is fine for systematic evaluation runs but painful for development.

The binary success metric hides capability progress. An agent that achieves arbitrary read/write but fails to bypass ASLR scores identically to one that crashes immediately. An agent that successfully escalates to root but miscalculates the flag file path gets zero credit. There's no partial scoring for "achieved primitive" or "defeated mitigation," which makes it hard to diagnose where agents are failing. Are they stuck on initial vulnerability triggering, primitive development, or exploit chaining? The logs might tell you, but the score won't.

Container images are static snapshots that will age poorly. A kernel task built against Linux 4.4 with 2016-era mitigations is already unrealistic—modern kernels have KASLR, SMAP, SMEP, and hardened usercopy. As exploitation techniques evolve (CFI, memory tagging, hardware-based isolation), the benchmark risks measuring historical capabilities rather than current offensive reality. The team versioned the benchmark (v1.0) and maintains a changelog, which suggests they're aware of this, but ongoing maintenance will be required.

Verdict

Use if: You're researching autonomous offensive security agents and need to measure end-to-end exploitation capability, not just vulnerability discovery. Use if you have institutional infrastructure to run multi-service Docker orchestration and can tolerate slow iteration cycles. Use if you need realistic constraints (network isolation, no internet, metered LLM calls) to prevent benchmark contamination. Use if you're publishing security research and need a credible, versioned benchmark that won't be dismissed as synthetic CTF toys.

Skip if: You're building vulnerability detection tools, fuzzers, or static analyzers—those need faster feedback loops than container orchestration provides. Skip if you're teaching exploitation fundamentals; pwn.college's generated challenges have better educational scaffolding. Skip if you lack Docker infrastructure or need to run evaluations on locked-down corporate networks. Skip if you need fine-grained capability metrics; binary pass/fail scoring won't help you debug agent failure modes. Skip if you're working on defensive AI applications—this benchmark is exclusively offensive.