> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

METATRON: Building a Penetration Testing Assistant with Local LLMs and Zero Cloud Dependencies

[ View on GitHub ]

METATRON: Building a Penetration Testing Assistant with Local LLMs and Zero Cloud Dependencies

Hook

Most AI-powered security tools send your reconnaissance data to OpenAI's servers. METATRON keeps everything local by treating Ollama as 'smart glue' between nmap, nikto, and a MariaDB audit trail—but the price of privacy is architectural fragility that makes it a teaching tool, not a weapon.

Context

Penetration testing has always been about tool orchestration—run nmap to find ports, feed results to nikto for web vulnerabilities, cross-reference CVEs, attempt exploits, document everything. The bottleneck isn't running tools; it's interpreting megabytes of raw output and deciding what to scan next. Commercial platforms like Metasploit Pro automate workflows but can't explain why a service looks vulnerable. GPT-4 can analyze security data brilliantly, but uploading client infrastructure data to OpenAI violates every responsible disclosure agreement ever written.

METATRON emerged from this gap: what if you could run a GPT-class model entirely on your Parrot OS laptop, feeding it tool output and getting conversational analysis without data ever leaving your machine? The project uses Ollama to run Qwen 3.5 locally, wrapping classic reconnaissance binaries in a Python orchestrator that creates an agentic feedback loop—the LLM requests targeted scans based on initial findings, stores everything in MariaDB, and generates PDF reports. It's deliberately simple: 1,200 lines of synchronous Python with no external APIs, no containers, no microservices. Just tools, a database, and a fine-tuned language model running in your terminal.

Technical Insight

Agentic Loop

shell=True subprocess

stdout text

raw scan output

scan results + context

HTTP POST

analysis + tool requests

response text

string matching parse

CRUD via raw SQL

star schema

history.sl_no FK

metatron.py CLI Loop

tools.py

llm.py Ollama Client

db.py MariaDB

System Binaries

nmap/nikto/whois/dig

Ollama API

Qwen 3.5 fine-tuned

MariaDB

vulnerabilities

fixes

exploits

System architecture — auto-generated

METATRON's architecture reveals how little code you need to build agentic workflows when you sacrifice robustness for directness. The core is four Python modules: tools.py shells out to system binaries, llm.py wraps Ollama's HTTP API, db.py handles MariaDB CRUD, and metatron.py orchestrates the scan loop. Here's the tool execution pattern:

def run_nmap(target, args="-sV -sC"):
    command = f"nmap {args} {target}"
    result = subprocess.run(command, shell=True, capture_output=True, text=True)
    return result.stdout

def run_nikto(target):
    command = f"nikto -h {target}"
    result = subprocess.run(command, shell=True, capture_output=True, text=True)
    return result.stdout

No timeouts, no sandboxing, no privilege separation—just raw subprocess.run() with shell=True. This is architecturally honest: METATRON assumes you're running on Parrot OS where these tools are pre-installed and you trust the LLM not to hallucinate malicious commands. The agentic loop works by parsing LLM responses for tool names:

def parse_ai_response(response_text, target):
    if "nmap" in response_text.lower():
        scan_result = run_nmap(target)
        return scan_result
    elif "nikto" in response_text.lower():
        scan_result = run_nikto(target)
        return scan_result
    # ... more string matching

This is the anti-pattern that makes METATRON fascinating. Instead of OpenAI-style function calling with JSON schemas ({"name": "run_nmap", "arguments": {"target": "..."}}), it just substring-matches natural language. If the model says 'Let's run nmap to check ports,' the code triggers nmap. Fragile? Absolutely. But it works with any LLM that can follow instructions, not just models fine-tuned for function calling.

The Modelfile approach is where things get clever. Instead of runtime prompt engineering, METATRON creates a derived Ollama model with baked-in pentesting context:

FROM qwen2.5:3b
PARAMETER temperature 0.7
PARAMETER top_k 10
PARAMETER top_p 0.9
PARAMETER num_ctx 16384
SYSTEM """
You are METATRON, a penetration testing assistant.
Analyze reconnaissance data and suggest next steps.
Request tools by name: nmap, nikto, whatweb, whois.
Provide CVE IDs when identifying vulnerabilities.
"""

This gets built once with ollama create metatron -f Modelfile, creating a persistent model that always operates in pentest mode. It treats model customization as infrastructure-as-code rather than per-request prompt injection. The tradeoff: you can't easily adjust behavior mid-scan without rebuilding the model.

Database design separates concerns operationally correctly. The exploits_attempted table is distinct from vulnerabilities:

CREATE TABLE vulnerabilities (
  vuln_id INT AUTO_INCREMENT PRIMARY KEY,
  history_id INT,
  vulnerability LONGTEXT,
  severity VARCHAR(20),
  cve_id VARCHAR(50)
);

CREATE TABLE exploits_attempted (
  exploit_id INT AUTO_INCREMENT PRIMARY KEY,
  history_id INT,
  exploit_name VARCHAR(255),
  payload LONGTEXT,
  success BOOLEAN
);

This matters for compliance—knowing you found a vulnerability is different from attempting to exploit it. Post-engagement reports need that separation. But notice the missing indexes on history_id foreign keys and the LONGTEXT fields with no size limits. After 100 scans with verbose nikto output, queries will crawl.

The DuckDuckGo integration is the hidden gem. Instead of paying for CVE database APIs, METATRON uses the duckduckgo-search library:

from duckduckgo_search import DDGS

def search_cve(cve_id):
    with DDGS() as ddgs:
        results = ddgs.text(f"{cve_id} exploit", max_results=5)
    return results

No rate limits, no authentication, no API costs. It's technically against some security aggregator terms of service, but for a local pentest tool, it's pragmatic. The limitation: DuckDuckGo results aren't structured CVE data—you get web snippets, not CVSS scores or patch timelines.

Gotcha

METATRON's simplicity is both its strength and fatal flaw. The two-terminal requirement—one running ollama serve, another running python metatron.py—exists because Ollama doesn't auto-start models on API requests. If you forget to preload the 8GB Qwen model into RAM, your scan hangs silently waiting for HTTP responses that never come. There's no startup check validating the model is running.

Security is theater-level bad. The default MariaDB password is '123', stored in plaintext in db.py. The tool that's supposed to find vulnerabilities ships with credentials a penetration tester would flag immediately. Worse, all reconnaissance runs with your user's full privileges—if you're root on Parrot OS (common for pentest distros), METATRON inherits that. A hallucinated command in the LLM response like rm -rf / would execute without question because subprocess.run(shell=True) has zero validation. The agentic loop that makes it smart also makes it dangerous if the model goes off-rails.

Portability is non-existent. Hardcoded assumptions about Parrot OS tool paths mean running this on Kali, Ubuntu, or macOS fails silently—tools aren't found, but the error handling just returns empty strings to the LLM, which then hallucinates analysis based on no data. The synchronous execution model means a 20-minute nikto scan blocks everything. No concurrency, no timeout handling, no way to scan multiple targets in parallel.

Verdict

Use if: You're learning AI agent architecture and want a readable codebase demonstrating local LLM integration without cloud dependencies, you're on Parrot OS and need a conversational interface to standard pentest tools for CTF competitions or educational labs, or you're prototyping custom security agents and want to fork a working Ollama/tool-chaining reference implementation. Skip if: You need production-grade red team tooling with audit trails that satisfy compliance requirements (the security issues and lack of sandboxing disqualify it immediately), you're scanning multiple targets or need concurrent execution (sequential blocking I/O makes this unusable at scale), you require robust error handling or cross-platform support (Parrot OS coupling and silent failures kill operational reliability), or you want mature LLM function calling with structured outputs (substring matching tool names is too brittle for real engagements). METATRON is a teaching tool that proves local LLMs can power security workflows—treat it as a reference architecture to learn from, not a platform to deploy.