Inside 0day-Rubbish: How LLM Ensembles Are Automating Vulnerability Discovery at Scale
Hook
Fifty-one zero-day vulnerabilities disclosed in six months—not by a nation-state APT group or a thousand-person security firm, but by an anonymous collective running LLM ensembles against obscure enterprise software. Welcome to industrialized exploit development.
Context
Traditional vulnerability research is artisanal work. A researcher spends weeks reversing a binary, fuzzing inputs, tracing execution flows, and handcrafting exploits. Then comes the 90-day coordinated disclosure dance: notify the vendor, negotiate timelines, wait for patches, maybe collect a bug bounty. The entire pipeline optimizes for vendor comfort over security outcomes.
0day-Rubbish flips this model. It's a disclosure-first repository where zero-days against enterprise software—HiveMQ brokers, GigaSpaces XAP clusters, Voicent IVR systems—appear bi-weekly with full proof-of-concept exploits, CVSS scores, and root cause analyses. No vendor coordination. No waiting periods. Just public GitHub commits with working RCE chains hours after discovery. The project claims its velocity comes from multi-LLM orchestration: Claude, GPT-4, DeepSeek, GLM, and Kimi working in concert to automate pattern recognition, static analysis, and exploit generation. It's vulnerability research reconceived as a content pipeline where AI handles discovery and humans just verify the output before hitting publish.
Technical Insight
The repository itself is architecturally unremarkable—a static file structure organized by vendor, version, and vulnerability class. What matters is the invisible orchestration layer: how five different LLMs collaborate to find exploitable bugs in software most researchers ignore.
The ensemble approach addresses individual model weaknesses. Claude excels at multi-step reasoning chains needed to trace data flow through complex business logic. GPT-4 provides broad pattern recognition across vulnerability classes. DeepSeek specializes in code-specific analysis with lower hallucination rates on syntax. GLM and Kimi bring extended context windows (200K+ tokens) crucial for analyzing sprawling enterprise codebases where a single transaction might touch fifteen different modules. The likely workflow: ingest target software documentation and decompiled code → each LLM performs independent analysis flagging suspicious patterns (unsanitized user input reaching eval(), XML parsers without entity restrictions, deserialization of untrusted data) → aggregate flagged locations → human researcher validates and develops PoC.
Consider a typical finding from Batch 5—CVE-2024-XXXX against KeyHelp hosting panel. The vulnerability is XML External Entity (XXE) injection in an API endpoint. Here's the conceptual LLM prompt chain that would surface this:
# Pseudocode for LLM ensemble XXE detection
targets = parse_repo("keyhelp-source") # Ingest codebase
# Stage 1: Pattern detection (DeepSeek optimized for code)
xml_parsers = deepseek.query(
context=targets,
prompt="""Identify all XML parsing operations. Flag any instances where:
1. Parser accepts external input without validation
2. DTD processing is enabled by default
3. External entity resolution isn't explicitly disabled"""
)
# Stage 2: Data flow analysis (Claude for reasoning chains)
entry_points = claude.query(
context=xml_parsers + targets,
prompt="""Trace user-controllable input from API endpoints to flagged XML parsers.
Map the complete data flow including any sanitization or validation layers."""
)
# Stage 3: Exploitability assessment (GPT-4 for breadth)
exploit_candidates = gpt4.query(
context=entry_points,
prompt="""For each identified flow:
- Confirm user input reaches parser unsanitized
- Verify external entity processing is enabled
- Assess impact (file read, SSRF, RCE via expect:// wrapper)
- Rate exploitability (authentication required? default configs vulnerable?)"""
)
# Stage 4: Long-context verification (Kimi for sprawling codebases)
full_chain = kimi.query(
context=targets, # Entire codebase as context
prompt=f"""Given exploit candidate: {exploit_candidates[0]}
Verify the complete attack chain including:
- Reachability from unauthenticated context
- Absence of WAF/input filters in default deployment
- Presence of sensitive files readable via XXE
Provide exploitation steps."""
)
# Human validation: researcher tests in lab environment
if verify_exploit(full_chain.exploitation_steps):
publish_disclosure()
The actual KeyHelp XXE exploit allows unauthenticated attackers to read arbitrary files via a crafted SOAP request—exactly the kind of vulnerability LLMs excel at finding because it's pattern-based (known dangerous API + missing security config) rather than requiring novel cryptographic or logic bug insights.
What makes this approach viable now is context window expansion. Analyzing a 50,000-line enterprise application was impossible for GPT-3's 4K tokens. Modern LLMs ingest entire repositories, maintaining awareness of how a user-supplied XML payload in api/v2/restore.php propagates through three middleware layers before reaching libxml_parse() with DTD processing enabled. The LLM doesn't need to understand business logic—just match patterns: "untrusted input" + "dangerous sink" + "no sanitization" = flag for human review.
The repository structure reinforces machine readability. Each disclosure follows a rigid schema:
vendor/
product-version/
vuln-type/
advisory.md # CVSS, timeline, remediation
exploit.py # Working PoC
analysis/
root-cause.md # Code-level explanation
traffic-capture.pcap
This isn't for human browsing—it's designed for automated threat intelligence ingestion. Purple teams can write scrapers that parse advisory.md for CVSS scores, extract exploit.py for red team playbooks, and feed traffic-capture.pcap into IDS signature generation. The standardization is the point: industrialize disclosure the same way LLMs industrialize discovery.
The limitation is validation overhead. LLMs hallucinate exploits constantly—flagging false positives where input validation exists three call stacks deep, or proposing exploitation paths blocked by default PHP configurations. The human researcher becomes a QA filter, testing each LLM-generated hypothesis in lab environments. This is still faster than manual review (LLMs screen thousands of potential issues overnight; humans validate tens), but it's not autonomous exploitation. The 'AI-driven' framing oversells automation when likely 70% of researcher time is still verification and PoC development.
Gotcha
The project's zero-day disclosure philosophy creates serious ethical and practical problems. Publishing exploits without vendor notification means patches don't exist when PoCs go public. For mainstream software, this creates brief windows of exposure. For the niche enterprise tools targeted here—Voicent IVR systems, Biamp conference hardware, Loadbalancer.org appliances—it creates months-long vulnerabilities because small vendors lack the resources for emergency patch cycles.
Defenders are put in impossible positions. A purple team discovers 0day-Rubbish published an RCE for their obscure ERP system. There's no patch. Vendor support says 'we're investigating.' Meanwhile the PoC is on GitHub with 120 stars and climbing. Your options: disable the vulnerable feature (breaking business processes), implement brittle WAF rules against the specific exploit (trivially bypassed with minor variations), or accept the risk. The repository provides zero defensive tooling—no IDS signatures, no hardening scripts, just a 'Recommended Mitigations' section that typically says 'upgrade to patched version' when no patched version exists.
There's also zero transparency into the discovery pipeline. The project publishes outputs (exploits) but not methods (LLM prompts, false positive rates, validation procedures). You can't assess whether findings are high-quality research or LLM hallucinations that happened to work once in a lab. The repository contains only eight PoCs despite claiming fifty-one cumulative disclosures—either earlier batches aren't committed or this is a showcase rather than a comprehensive archive. For a project positioning itself as advancing the field, the opacity around methodology is disqualifying.
Verdict
Use if: You're a red teamer targeting specific enterprise stacks where vendor patching is glacial and you need working exploits against obscure software (KeyHelp hosting panels, HiveMQ brokers, GigaSpaces XAP clusters). This delivers curated 1days 48 hours after disclosure—high-quality PoCs you won't find elsewhere. Also valuable as a case study in LLM-assisted vulnerability research workflows if you're building similar tooling and want to understand ensemble approaches to static analysis. Skip if: You need responsible disclosure timelines, defensive intelligence, or ethical vulnerability research. This is full-disclosure maximalism that creates risk for defenders without providing tools to mitigate it. Also skip if you want reusable code—there's no framework to fork here, just a publication pipeline. For most practitioners, traditional alternatives (Nuclei templates with vendor coordination, Metasploit's vetted modules, or ZDI's coordinated disclosures) provide better risk-reward ratios. This is provocative as philosophy but thin as infrastructure.