> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

ReAgent: How Dual-Agent LLMs Reconstruct Source Code from Binaries Without Hallucinating

[ View on GitHub ]

ReAgent: How Dual-Agent LLMs Reconstruct Source Code from Binaries Without Hallucinating

Hook

Upload decompiled code to ChatGPT and you'll get plausible-looking C++ that compiles but computes the wrong values 30% of the time. ReAgent solves this with architectural skepticism: one LLM writes code, another tries to reject it.

Context

Reverse engineering compiled binaries into readable source code has traditionally been a grinding manual process. Tools like Ghidra and IDA Pro produce decompiled pseudocode that's technically correct but littered with variable names like uVar3 and control flow that looks like it was written by a compiler having a bad day. For game preservation projects reconstructing titles like GTA or OpenRCT2, security researchers analyzing malware, or anyone maintaining legacy systems where the original source was lost, the goal isn't just understanding what the binary does—it's producing buildable, maintainable C/C++ that integrates into a modern codebase.

The obvious move in 2024 is feeding decompiler output to an LLM and asking it to clean things up. This works surprisingly well for simple functions, but falls apart at scale. LLMs are confident confabulators: they'll rename variables sensibly, add helpful comments, and restructure loops into readable patterns while silently changing the algorithm. A security researcher analyzing a crypto function might get code that looks cleaner but implements a different cipher. A game preservation team might get physics code that compiles and almost works, failing only in edge cases that won't surface until months later. The fundamental problem is that LLMs have no built-in mechanism for uncertainty—they optimize for plausibility, not correctness. ReAgent's architecture directly confronts this by treating LLM output as adversarial and implementing four independent validation gates that code must pass before it's considered acceptable.

Technical Insight

ReAgent's core innovation is architectural distrust. Instead of a single LLM generating code and self-validating it (which consistently produces overconfident garbage), it implements a dual-agent system where a 'reverser' agent proposes C/C++ implementations and a separate 'checker' agent evaluates them with no visibility into the reverser's reasoning. This isn't just prompt engineering—it's a genuine information barrier. The checker receives only the candidate code and structural constraints derived from the binary, preventing the confirmation bias that plagues single-agent systems.

The evidence pipeline feeds both agents through a read-only Ghidra backend. ReAgent extracts decompilation output, cross-references, control flow graphs, and normalized P-code (Ghidra's intermediate representation) into structured JSON bundles. Critically, the LLM cannot modify analysis state or request arbitrary Ghidra operations—it works with preloaded evidence plus a bounded set of investigation operations (vtable lookups, string references, global variable analysis) requested explicitly through an investigation round mechanism. This prevents the context window explosion that kills most LLM-driven analysis tools while still enabling adaptive evidence gathering:

# Simplified investigation round structure from ReAgent's orchestrator
investigation_results = []
for requested_op in reverser_agent.investigation_requests:
    if requested_op.type == "vtable_lookup":
        vtable_data = ghidra_backend.get_vtable(
            requested_op.class_address,
            max_entries=50  # Hard limit prevents unbounded context
        )
        investigation_results.append(vtable_data)
    elif requested_op.type == "xrefs_to":
        xrefs = ghidra_backend.get_xrefs(
            requested_op.target_address,
            max_depth=2  # Prevents reference chain explosion
        )
        investigation_results.append(xrefs)

# Investigation results are appended to context, but original
# evidence bundle remains immutable
reverser_context = evidence_bundle + investigation_results

Once the reverser produces a candidate, validation runs through four independent gates. The first is the checker agent itself, which performs semantic review: Does the code match the control flow structure? Are memory operations consistent with the decompiled patterns? Does variable lifetime match register usage? The second gate is structural verification—an automated AST/CFG comparison between the candidate's compiled form and the original binary's normalized P-code. This is smarter than raw assembly diff because it abstracts register allocation and instruction selection while preserving data dependencies:

# P-code normalization catches semantic differences decompilers hide
original_pcode = normalize_pcode(
    ghidra_backend.get_pcode(original_function),
    strip_registers=True,  # Abstract register allocation
    normalize_constants=True  # 0x1 and 1 are equivalent
)
candidate_pcode = normalize_pcode(
    compile_and_extract_pcode(candidate_code),
    strip_registers=True,
    normalize_constants=True
)
similarity_score = compare_cfg_patterns(original_pcode, candidate_pcode)
if similarity_score < 0.85:  # Tunable threshold
    return ValidationFailure("P-code structure mismatch")

The third gate is build-test-runtime validation, and it's where ReAgent's trust model gets interesting. The system can execute arbitrary build commands (make test, cargo test, project-specific scripts) in copy-on-write project snapshots—symlink-based workspace copies that preserve the original source tree. But here's the honest limitation: ReAgent cannot verify that make test actually validates the candidate function. A test suite might have poor coverage, or tests might pass despite semantic differences if the function isn't exercised properly. So ReAgent requires explicit human attestation via a trust_configured_commands flag. This isn't a cop-out—it's admitting that shell command execution isn't semantic validation, forcing users to consciously decide whether their test infrastructure is meaningful.

The fourth gate is parity analysis: 11 heuristic signals that detect code smells the other gates might miss. Does the candidate contain TODO markers or stub implementations? Is there excessive plugin-call density suggesting the LLM over-relied on helper functions? Are there floating-point sensitivity mismatches (original binary uses doubles, candidate uses floats)? These are separated into RED signals (blocking failures like stub markers) and YELLOW signals (advisory warnings like call-count mismatches). This two-tier system makes ReAgent useful for exploration—you can review YELLOW-flagged code manually—without pretending heuristics are proofs.

The orchestrator maintains session state in atomically-written JSON, preserving round diagnostics and candidate history. This enables resume-from-checkpoint workflows critical for expensive multi-hour reverse engineering sessions where you're burning through Claude API tokens at scale. The cumulative validation phase can compose accepted functions into scratch projects for class-level integration testing, catching interface mismatches that per-function validation misses.

Gotcha

ReAgent's documentation repeatedly emphasizes 'conservative verification,' and you need to internalize what this means: passing all four validation gates does NOT prove semantic equivalence. A candidate can have identical control flow, pass structural P-code comparison, compile and pass tests, and clear parity analysis while still containing logic bugs invisible to these heuristics. Complex conditionals, floating-point edge cases, undefined behavior exploitation—these can all slip through. You're still manually auditing generated code for correctness; ReAgent just gives you much better tooling than 'vibes-based ChatGPT review.'

The build validation gate's trust model exposes a real limitation: if your project lacks meaningful tests, or if tests exist but have poor coverage of the functions you're reversing, the third validation gate becomes security theater. ReAgent can't auto-generate tests that prove correctness, and the trust_configured_commands flag is an explicit admission that humans must verify test quality. For projects without existing CI infrastructure—like most malware analysis scenarios or legacy codebases where tests were never written—you're relying heavily on the structural and parity gates, which are purely heuristic. The tool also creates no dependency graph between functions, so re-reversing a function after finding a bug doesn't automatically invalidate dependent code that called the old implementation. Large refactoring efforts require manual coordination to track which functions need re-validation.

Verdict

Use if: You're working on large-scale reverse engineering projects (game preservation, legacy system reconstruction) where you need buildable, integration-tested C++ and you have existing test infrastructure that actually exercises the code. The dual-agent architecture and multi-gate validation provide genuine rigor for batch reconstruction of thousands of functions, and the investigation loop prevents context explosion while enabling adaptive evidence gathering. The copy-on-write validation and session resume make it practical for expensive multi-day reversing efforts. Skip if: You're doing exploratory analysis where Ghidra's decompiler output is sufficient for understanding, you lack meaningful tests (the validation becomes theater), or you're working on codebases too complex for LLM-generated code to reliably pass structural verification. Also skip if you're expecting semantic equivalence proofs—ReAgent is honest about providing 'conservative verification,' which means you're still auditing generated code, just with better tooling than raw LLM output.