> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

The AI Security Benchmark Explosion: Navigating 175 Evaluation Frameworks

[ View on GitHub ]

The AI Security Benchmark Explosion: Navigating 175 Evaluation Frameworks

Hook

There are now more AI security benchmarks than there are days in a year—175 at last count—and almost none of them agree on what 'secure' means or how to measure it.

Context

Two years ago, evaluating whether an LLM was 'secure' meant running HumanEval and hoping it didn't suggest SQL injection. Today, security researchers face an embarrassment of riches: CTF environments with Docker sandboxes, repository-level vulnerability datasets, real-world bug bounty challenges with dollar values attached, and knowledge benchmarks testing everything from cryptography theory to OWASP Top 10 awareness.

The Awesome-AI-Security-Benchmarks repository emerged from this chaos as a single-page index to the entire ecosystem. Maintained by individual contributor Evan Thomas Luke, it catalogs benchmarks across five major categories: knowledge/Q&A tests, capture-the-flag environments, penetration testing harnesses, vulnerability detection datasets, and secure code generation frameworks. Unlike Papers With Code or academic surveys, this list specifically targets practitioners who need to answer: 'Which evaluation should I run before deploying this code-generation model?' The repository doesn't provide tooling or analysis—it's purely a reference index with links to papers, datasets, and GitHub repos.

Technical Insight

Browse Categories

API Calls

Docker Execution

Agent Interaction

Static Analysis

Code Review

Manual Updates

Scraping

Security Researchers

ML Engineers

Taxonomy Layer

6 Evaluation Types

Static Knowledge

Q&A Datasets

Interactive CTF

Environments

Penetration Testing

Harnesses

Vulnerability Detection

Datasets

Secure Code

Generation

External Resources

Papers/Datasets/Repos

Community

Contributors

Academic Papers

arXiv/IEEE/ACM

System architecture — auto-generated

The repository's architecture is deliberately minimal: a single README.md file containing markdown tables organized by evaluation category. Each entry follows a consistent schema:

| Benchmark Name | Description | Link | Year |
|----------------|-------------|------|------|
| CyberSecEval 2 | Tests for insecure code, prompt injection | [Paper](url) | 2024 |

What makes this useful isn't technical sophistication—it's the taxonomy. The list distinguishes between three fundamental evaluation paradigms that are often conflated in security ML papers:

Static Knowledge Tests (SecEval, CyberMetric, SECURE): These are multiple-choice or short-answer datasets testing whether models understand security concepts. Running them requires nothing more than API calls:

# Typical static benchmark usage
from datasets import load_dataset

eval_data = load_dataset("cybermetric/cybersecurity-qa")
for question in eval_data:
    response = model.generate(question['prompt'])
    score = check_answer(response, question['correct_answer'])

These benchmarks are cheap to run but measure recall, not capability. A model can ace a cryptography quiz while still generating timing-attack-vulnerable comparison functions.

Interactive Environments (NYU CTF Bench, Cybench, InterCode-CTF): These require sandboxed execution environments—usually Docker containers with vulnerable applications or CTF challenges. The model must interact with running systems:

# CTF benchmark pattern (simplified from InterCode)
from intercode import CTFEnvironment

env = CTFEnvironment(challenge="buffer_overflow_basic")
state = env.reset()

for attempt in range(max_attempts):
    action = agent.get_action(state)  # Model generates shell commands
    state, reward, done = env.step(action)
    if done and reward > 0:
        # Flag captured
        break

These benchmarks expose whether models can execute attacks, not just describe them. The NYU CTF Bench contains 200+ challenges ranging from basic web exploits to binary reverse engineering. The infrastructure cost is real—Cybench documents requiring GPU instances and hourly AWS charges for their sandboxed eval harness.

Repository-Level Analysis (SecRepoBench, A.S.E, emerging in 2025): The newest category reflects a critical shift. Early vulnerability datasets like Juliet Test Suite provided isolated C functions with buffer overflows. Modern benchmarks like SecRepoBench provide entire codebases where vulnerabilities span multiple files:

# Repository-level vulnerability detection
from secrepobench import load_repo

repo = load_repo("web-app-2024-001")  # Full Django project
vulns = model.analyze_repository(
    repo.source_tree,
    context_window=100000,  # Needs long-context models
    analysis_type="taint_tracking"
)

# Vulnerability might involve:
# - User input in views.py
# - Sanitization logic in utils.py  
# - Database query in models.py

This matches real-world security review workflows where understanding data flow across architectural boundaries matters more than spotting strcpy() calls.

Digging into specific benchmarks reveals important patterns. Meta's CyberSecEval series (v1 through v4) shows rapid threat model evolution: v1 focused on 'will it write buffer overflows,' v2 added prompt injection resistance, v3 introduced automated exploit generation, and v4 (2024) now tests SOC analyst workflows like alert triage and incident response. Each version deprecates previous metrics as attacks evolve.

The CTF cluster—InterCode-CTF, 3CB, CTF-Dojo—consistently hovers around 40-200 challenges, suggesting task curation is the limiting factor. These benchmarks require security experts to manually design challenges, validate solutions, and maintain infrastructure. Contrast this with BountyBench's approach: scrape real HackerOne/Bugcrowd reports, containerize the vulnerable versions, and grade based on whether the model finds the actual CVE. Economic signal (actual bounty paid) becomes ground truth.

A concerning pattern emerges in the vulnerability detection category: heavy bias toward C/C++ memory safety. Juliet Test Suite, PrimeVul, ARVO, SV-TrustEval-C all focus on buffer overflows, use-after-free, and integer overflows. Meanwhile, cloud-native attack surfaces—IAM misconfigurations, SSRF in microservices, GraphQL injection—barely appear. The benchmark landscape mirrors what academic security researchers publish on, not what causes breaches.

Gotcha

The repository's biggest limitation is what it doesn't tell you. When comparing five different CTF benchmarks, you get no guidance on difficulty levels, contamination risks, or reproducibility track records. Did GPT-4's training data include solutions to InterCode-CTF challenges posted on GitHub? The list won't tell you. Several benchmarks link to papers behind ACM paywalls with no accessible dataset links, making them discovery-only citations rather than usable tools.

Maintenance is a single point of failure. The repository is one person's manual curation effort with explicit warnings about 'errors and hallucinations.' Links rot—several entries point to arXiv preprints that were later accepted to conferences with different final URLs. There's no CI/CD checking link validity, no community review process, and no guarantee the descriptions match current benchmark versions. The Discord link suggests community engagement, but the contribution process is opaque.

Categorization errors appear throughout. Some entries mix datasets (static vulnerability examples) with frameworks (full evaluation harnesses) and platforms (CTF hosting infrastructure) in the same table. The 'Secure Code Generation' category contains both training datasets for fine-tuning and evaluation benchmarks for testing, which serve fundamentally different purposes. For practitioners trying to build eval pipelines, these distinctions matter enormously.

Verdict

Use if you're a security researcher writing related work sections, an ML engineer pitching evaluation infrastructure to leadership (the sheer volume makes the case), or a practitioner doing initial discovery before selecting benchmarks for a specific threat model. The breadth is genuinely valuable—finding that BountyBench exists or discovering SecRepoBench's repository-level approach could redirect an entire eval strategy. Treat it as a research index, not a review guide. Skip if you need comparative analysis, reproducibility guarantees, or production-ready tooling. The repository provides zero guidance on which benchmarks actually measure what they claim, which have contaminated training data, or which are maintained versus abandoned. If you're building eval infrastructure today, mine this for ideas but immediately move to the primary sources—read the papers, check GitHub activity, and verify datasets are accessible. For runnable harnesses, go directly to Meta's CyberSecEval or BountyBench rather than treating this list as a menu to choose from.