Cairn: When Autonomous Agents Coordinate Through Shared Reality Instead of Chat
Hook
The only AI framework to win a major security competition solved all 54 challenges using agents that had never spoken to each other—they communicated exclusively by reading and writing to a shared graph.
Context
Multi-agent AI systems typically fail in predictable ways: agents spam each other with redundant information, lose track of who knows what, and spend most of their context windows recapping previous conversations. This 'prompt stuffing' problem gets worse as problems grow complex—by the tenth message in a chain, half the tokens are just catching everyone up. The underlying issue is architectural: we've bolted a chat interface onto systems that need coordination, not conversation.
Cairn emerged from penetration testing research, where autonomous exploration hits this wall immediately. You can't script pentesting workflows—every target has unique attack surfaces, dead ends, and pivot opportunities. Existing tools like Metasploit codify known exploits but require human operators to decide what to try next. Meanwhile, agentic frameworks like AutoGPT thrash aimlessly because they lack domain grounding. The breakthrough in Cairn is treating security assessment as state-space search with a blackboard architecture: agents never communicate directly, instead reading and writing to a shared graph of facts (confirmed findings) and intents (hypotheses to explore). This maps perfectly to how human pentesters work—you document what you've found, note what looks promising, and let anyone on the team pick up the next thread.
Technical Insight
Cairn's architecture inverts the typical multi-agent design. Instead of spawning specialized agents (a 'reconnaissance agent', 'exploitation agent', etc.) with fixed roles, it maintains a central graph and generates tasks dynamically based on topology. The dispatcher reads the graph, identifies unexplored intents, and spawns ephemeral worker containers that execute an OODA loop: observe the entire graph state, orient to their position relative to the goal, decide on new intents, and act by exploring then writing results back as facts.
The graph structure is deceptively simple. Facts are confirmed discoveries with evidence—an open port, a discovered credential, a successful command execution. Intents are directions to explore, spawned from facts ('this service version has known CVEs' → intent to attempt exploit). Hints are human guidance injected into the graph. Workers are model-agnostic: they wrap Claude, GPT-4, or local LLMs, receive structured prompts synthesized from graph context, and return JSON.
Here's what a worker observes when waking up:
# Simplified graph state passed to worker
graph_state = {
"facts": [
{
"id": "fact_1",
"content": "Port 22 open, SSH banner: OpenSSH_7.4",
"evidence": "nmap scan output",
"spawned_intents": ["intent_3"]
},
{
"id": "fact_2",
"content": "Web server at :8080 returns 401 Unauthorized",
"evidence": "curl response headers"
}
],
"intents": [
{
"id": "intent_3",
"parent_fact": "fact_1",
"hypothesis": "OpenSSH 7.4 vulnerable to user enumeration",
"status": "pending"
},
{
"id": "intent_4",
"parent_fact": "fact_2",
"hypothesis": "Web server may have default credentials",
"status": "pending"
}
],
"goal": "Obtain root access and read /flag.txt"
}
The worker receives this entire context, picks an unexplored intent, attempts it, then writes results back. Critically, workers never see each other's intermediate reasoning—only the facts they've committed to the graph. This stigmergic coordination (communication through environmental modification, borrowed from ant colony behavior) prevents the information silos that plague conversational agents.
The three-task taxonomy reveals domain grounding. Bootstrap tasks attempt to solve the challenge immediately with zero exploration—grabbing low-hanging fruit before expensive graph traversal. Reason tasks generate new intents from existing facts without execution (strategic planning). Explore tasks execute specific intents and produce new facts (tactical action). The dispatcher schedules these based on graph topology, not hardcoded workflows.
Local execution mode is architecturally significant. Instead of containerizing workers, it runs them directly on the host with your pre-authenticated CLI tools:
# Local mode bypasses Docker entirely
if config.execution_mode == "local":
worker_env = os.environ.copy() # Inherit host auth
worker_process = subprocess.Popen(
["python", "worker.py", "--graph", graph_path],
env=worker_env,
cwd=project_dir
)
This sidesteps credential management and lets agents use native tools like aws, kubectl, or gcloud with your existing authentication. The security trade-off is obvious—an LLM now has your full user privileges—but for offensive security research, this is often acceptable.
The dispatcher's role as sole protocol writer ensures linearizability. Workers submit facts and intents as proposals; the dispatcher serializes writes to the graph. This prevents race conditions where two workers explore the same intent redundantly, or write contradictory facts that create inconsistent worldviews. The graph becomes the single source of truth, and workers trust it completely.
What makes this work for penetration testing specifically is the fact-intent duality mapping to exploit development's natural rhythm. Facts are foothold states (confirmed access, discovered assets). Intents are attack hypotheses (this might be exploitable). Human pentesters already think this way—Cairn just formalizes it into a coordination primitive that agents can execute autonomously. The competition validation (54/54 challenges solved) suggests this mapping generalizes: anywhere you have black-box exploration toward a defined goal state, this architecture applies.
Gotcha
The codebase has no graph pruning strategy. As exploration continues, the graph accumulates dead-end intents and obsolete facts. Workers observe the entire graph on every wake-up, so context windows fill with noise. There's no cost model, no prioritization heuristic, no mechanism to mark intents as 'exhausted' or deprecate facts when the state space shifts. On complex targets with deep attack graphs, you'll hit context limits and waste tokens re-exploring known dead ends.
Security surface is massive. Default Docker mode gives worker containers full socket access—if an LLM is manipulated or hallucinates hostile commands, it can escape the container trivially. Local mode is worse: agents inherit your full user privileges with zero sandbox. The repository includes no prompt injection defenses, no command validation, no audit logging. This is acceptable for CTF competitions in isolated environments, but deploying this against production systems or with untrusted LLMs is asking for disaster. There's also zero reproducibility tooling—no way to replay a graph evolution, no metrics beyond solve/no-solve, no A/B testing infrastructure for comparing prompting strategies. You run it, it either solves the challenge or burns your token budget, and you have limited insight into why.
Verdict
Use if: You're doing offensive security research or CTF competitions where the goal state is unambiguous ('read /flag.txt'), the environment is black-box, and you value autonomous exploration over predictable execution. The stigmergic coordination through facts and intents is the cleanest solution I've seen for multi-agent systems that need to avoid prompt-stuffing and information silos, and the architecture is genuinely domain-agnostic—swap the prompts and goal state, and this works for any state-space search problem. Skip if: You need cost controls, audit trails, or deterministic behavior. The lack of graph pruning means token costs spiral on complex targets, the security surface is unacceptable for production environments, and the system will happily explore expensive dead ends with no way to constrain search. If you're operating under compliance requirements or need to explain findings to clients, stick with Metasploit and manual pentesting—Cairn is for researchers who want emergence and can tolerate chaos.