> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

The-Scaffolding: Building Multi-Agent Orchestration for Offensive Security

[ View on GitHub ]

The-Scaffolding: Building Multi-Agent Orchestration for Offensive Security

Hook

Most AI security agents fail after 20 minutes because they forget what they scanned, repeat the same reconnaissance, and hallucinate vulnerabilities from raw tool output. The-Scaffolding treats these as architectural constraints, not prompt engineering problems.

Context

The offensive security community has been experimenting with LLM-powered agents for CTF challenges and bug bounty hunting since GPT-4's release, but production deployments remain rare. The typical failure pattern is predictable: agents burn through context windows re-running nmap scans, misinterpret Burp Suite output, lose track of which directories they've enumerated, and generate false positive reports that waste analyst time. Frameworks like AutoGPT and LangChain offer general-purpose agent capabilities—memory systems, planning loops, tool integration—but they lack domain-specific knowledge about attack chains, confidence scoring for security findings, or benchmarks that measure true positive rates on actual vulnerabilities.

The-Scaffolding emerged from this gap as a meta-orchestration layer specifically designed for security workflows. Instead of building yet another agent runtime, it provides infrastructure that sits between existing AI agents (Claude, GPT-4, Cursor) and offensive security tooling (Metasploit, sqlmap, Nessus). The core thesis is that reliable security agents require three components absent from generic frameworks: externalized state management to survive long reconnaissance sessions, pre-filtered tool output with confidence scores to reduce hallucinations, and reproducible benchmarks that measure exploit success rates rather than conversation quality. Built around the Model Context Protocol (MCP), The-Scaffolding implements this as a plug-and-play harness where you bring your own agent and the scaffold handles tool access, skill loading, and performance validation.

Technical Insight

MCP Servers

Initialize Project

Generate Configs

.mcp.json

.codex/config.toml

env vars

Load by Challenge Type

Load by Challenge Type

Tool Requests

Tool Requests

Confidence-Scored Findings

Execution Results

PoC Code

Solve Outcomes

Validate

Compare MCP vs Skills-Only

scaffold.py CLI

router-spec.json

Multi-Agent Configs

Claude Desktop

Cursor/Codex

Other Agents

Skills Library

Markdown Files

MCP Server Layer

LatticeMind

Scanning + Scoring

Kali MCP

Security Tools

GitHub MCP

Exploit Discovery

External Notes

State Persistence

Benchmark System

System architecture — auto-generated

The architecture follows a hub-and-spoke pattern with the scaffold.py CLI as the orchestrator. When you initialize a project, it generates a multi-agent configuration from a single router-spec.json, creating .mcp.json for Claude Desktop, .codex/config.toml for Cursor, and equivalent configs for six supported agents. This prevents the configuration drift that kills multi-client integrations—change your MCP server URLs once, regenerate, and all agents inherit the new setup.

The skill system is where domain knowledge lives. Each skill is a markdown file encoding recon patterns, attack chains, and validation rules for specific challenge types (SQLi, XSS, privilege escalation). Here's a simplified example of what a skill contract looks like:

# In skills/web-sqli-basics.md
## Target: SQL Injection Detection
## Baseline Metrics:
# - Phase Router Success: 85%
# - True Positive Rate: 0.78
# - False Positive Rate: 0.12

### Reconnaissance Pattern
1. Identify user input vectors (GET/POST params, headers, cookies)
2. Test for error-based SQLi with single quote injection
3. Validate with boolean-based blind payloads if errors suppressed
4. Confirm exploitability with time-based techniques

### Validation Rules
- HTTP 500 + database error message = High confidence finding
- Differential response times (>5s delay) = Medium confidence
- Content-length changes without errors = Low confidence, requires manual review

### Expected Outputs
confidence_score >= 0.7 for reporting
vulnerability_type: "SQL Injection"
affected_parameters: [list]
proof_of_concept: [payload that triggered detection]

Agents selectively load relevant skills based on challenge type to avoid context bloat. For a web CTF, only web-sqli-basics.md, xss-detection.md, and directory-traversal.md get loaded—not the binary exploitation or network pivoting skills. The feedback loop writes solve outcomes back: if an agent successfully exploits a blind SQLi that the skill rated low-confidence, the baseline metrics update to reflect that the validation rules were too conservative.

Tool access happens through three MCP servers with distinct responsibilities. LatticeMind provides automated scanning with confidence-scored findings—instead of agents parsing raw sqlmap output (which causes hallucinations), they receive structured results: {"vulnerability": "SQLi", "confidence": 0.89, "location": "/api/user?id="}. Kali MCP wraps deterministic security tools (nmap, gobuster, ffuf) for agent execution. GitHub MCP enables exploit and proof-of-concept discovery from public repositories. This three-layer hierarchy separates automated recon (high latency, probabilistic), tool execution (low latency, deterministic), and knowledge retrieval (medium latency, contextual).

State persistence is externalized into a notes system rather than relying on agent memory or expanding context windows. After each reconnaissance phase, agents write findings to structured notes:

{
  "timestamp": "2024-01-15T14:32:11Z",
  "phase": "network_enumeration",
  "findings": [
    {
      "service": "Apache 2.4.49",
      "port": 80,
      "confidence": 0.95,
      "exploitability": "CVE-2021-41773 path traversal",
      "attempted": false
    }
  ],
  "next_actions": ["test_path_traversal", "enumerate_directories"]
}

On subsequent runs, agents resume from these notes rather than re-scanning. This pattern is critical for multi-hour CTF challenges where context window costs would otherwise become prohibitive.

Benchmarking uses a multi-tiered approach targeting different validation concerns. Phase-routing tests run cheap sanity checks: given a challenge description, does the agent correctly identify it as web/binary/crypto and load appropriate skills? Security benchmarks compare five profiles (control with no tools, skills-only, adaptive, MCP-enabled, skills-lite) on vulnerable test applications, measuring true positive rates, false positives, and false negatives with ground-truth labels. The CTF testbench deploys five dockerized challenges requiring full exploitation chains and mandatory writeups graded manually—this catches the cases where agents report success but didn't actually capture flags. Each tier trades off cost and signal: phase routing costs pennies per run for basic validation, while full CTF runs cost dollars but prove end-to-end capability.

Gotcha

The MCP server dependency creates a single point of failure. If LatticeMind is down, returns low-confidence noise, or hits rate limits, the entire harness degrades. There's no visible fallback strategy—agents don't gracefully degrade to raw tool execution or switch to alternative scanners. In production bug bounty work, this means you'd need redundant scanning infrastructure and retry logic that the current codebase doesn't provide.

The self-improving skill loop is conceptually interesting but practically risky. When agents write solve outcomes back into skill files, there's no versioning, conflict resolution, or automated rollback. A single misinterpreted success could corrupt validation rules: if an agent reports exploiting SQLi when it actually just found a benign error message, the skill's false positive rate gets baked in. The validation scripts catch schema violations but not subtle semantic errors. For serious use, you'd need to fork the skill files, implement version control, and manually review updates before merging them back—which defeats the automation benefit. The project also shows signs of early-stage maturity: three GitHub stars, no community contributions, no evidence of testing against real-world platforms like HackTheBox or production bug bounty programs. The benchmarks run against dockerized test challenges, which is fine for validation but doesn't prove the harness survives the chaos of actual offensive security work—WAFs that poison responses, rate limiting, multi-step authentication, ephemeral infrastructure.

Verdict

Use if: You're an AI engineer or security researcher building agent-driven offensive workflows and you're already hitting context limits with existing frameworks. The-Scaffolding is a reference implementation that demonstrates externalized state management, confidence-scored findings, and multi-agent orchestration patterns worth adopting. The benchmarking infrastructure alone—particularly the multi-tiered approach with objective metrics—is more rigorous than most academic security AI projects and provides a template for evaluating your own systems. Fork it, instrument the MCP layer for observability, version the skill system properly, and extend the architecture with your domain knowledge. Skip if: You need production-ready tooling for operational bug bounty or red team engagements. The MCP dependency creates fragility, the three-star GitHub presence signals unproven reliability, and you'll spend more time debugging the harness than solving challenges. Most working offensive security professionals should stick with custom Python automation (requests, pwntools, scapy) + Claude API for full control, or use established frameworks like Metasploit with LLM-guided tool selection. Also skip if you're evaluating AI security tools for immediate deployment—this is research-grade infrastructure that assumes you'll modify it extensively, not a turnkey solution.