Numbat: Stopping AI Agents Before They Exfiltrate Your Secrets
Hook
Your AI coding assistant can read your .env file, extract AWS credentials, and POST them to an external API—all before your EDR notices. Numbat catches it at the moment of decision, not after the damage.
Context
AI coding agents like Claude Code, GitHub Copilot Workspace, and OpenAI's Codex are crossing the Rubicon from autocomplete suggestions to autonomous tool execution. They create files, run shell commands, browse documentation, and commit code without waiting for your approval on every action. This shift from advisory to executive power creates a security gap that existing tools can't fill.
Traditional endpoint detection responds after processes execute—your EDR sees a file write or network connection, correlates it with threat intelligence, and alerts. But by then, credentials are already in flight. Application-layer tracing like LangSmith captures LLM API calls but misses what happens client-side when the agent's tool invocation actually executes. Process monitoring via eBPF sees syscalls but can't distinguish 'developer opened .aws/credentials' from 'agent opened .aws/credentials' because they're the same process. Numbat, built by Perplexity's security team to protect their own agent deployments, solves this by hooking into agent runtimes at the exact moment an AI proposes an action—before the tool call executes—and running detection rules that can block high-risk operations synchronously.
Technical Insight
Numbat's architecture hinges on a clever separation of concerns: vendor-specific instrumentation that speaks each agent's native protocol, a normalization layer that maps chaos into canonical events, and a pure-Go CEL engine that evaluates rules against both atomic events and multi-step sequences. The instrumentation layer isn't monolithic—it adapts to what each agent ecosystem supports. For Claude Desktop and Cline, it implements MCP (Model Context Protocol) servers that intercept tool calls. For OpenAI Codex, it uses Codex-specific hooks. For platforms without live instrumentation, it falls back to forensic parsing of on-disk session artifacts.
The pre-action hook is the killer feature. Here's a simplified example of how Numbat intercepts a file creation before it happens:
// Pseudo-code representation of pre-action flow
type ToolCallEvent struct {
SessionID string
ToolName string
Arguments map[string]interface{}
Source string // "claude-code", "codex", etc.
}
func (h *PreActionHook) Evaluate(event ToolCallEvent) EnforcementDecision {
// Normalize to canonical schema
canonical := normalizeEvent(event)
// Run CEL rules against normalized event
for _, rule := range h.rules {
if rule.Enforce && rule.CELExpr.Eval(canonical) {
return EnforcementDecision{
Action: "deny",
Reason: rule.Message,
RuleID: rule.ID,
}
}
}
return EnforcementDecision{Action: "allow"}
}
When Claude Code calls create_file with path .env, Numbat intercepts it before the filesystem write, evaluates CEL rules like tool.name == 'create_file' && args.path.matches('.*\\.(env|secret|key)$'), and returns a deny decision that Claude respects by showing an error instead of executing. The timing is everything—this happens in the decision loop, not the cleanup phase.
The normalization layer is where Numbat's multi-agent ambitions shine. Each agent uses different schemas: Claude Code emits {"tool": "create_file", "path": "..."}, OpenClaw sends before_tool_call webhooks with nested JSON, Codex uses its own tool schema. Numbat maps all of these into a canonical event model:
{
"version": "0.2.0",
"type": "event",
"timestamp": "2024-01-15T10:30:00Z",
"session_id": "sess_abc123",
"tool_call_id": "call_xyz789",
"event_type": "tool_call",
"tool": {
"name": "create_file",
"args": {
"path": ".env",
"content": "AWS_SECRET_KEY=..."
}
},
"source": {
"agent": "claude-code",
"user": "dev@company.com"
}
}
This canonical format lets you write one rule that works across all supported agents. The preserved source references mean forensic analysis can trace back to vendor-specific logs when investigating incidents.
Stateful sequence detection handles multi-step attacks without a database. Numbat holds a fixed-size window of recent events per session in memory. A rule like 'agent read secrets then made network request' works by evaluating against the event stream:
// CEL expression for sequence detection
events.exists(e, e.tool.name == 'read_file' &&
e.tool.args.path.contains('secret')) &&
events.exists(e, e.tool.name == 'http_request' &&
e.timestamp > events.filter(x, x.tool.name == 'read_file')[0].timestamp)
The window is scoped by session_id, so sequences don't leak across unrelated agent runs. When the window fills, old events drop out—no persistence, no database, just a bounded in-memory structure that makes the detector truly endpoint-local. This design choice trades perfect recall for operational simplicity: you can't query 'show me all sessions from last month,' but you also don't need Postgres running on every developer laptop.
The rule system itself uses CEL (Common Expression Language), the same expression language Kubernetes uses for validation. Rules live in YAML files with versioning:
version: 0.2.0
rules:
- id: block-ssh-key-write
name: Block SSH private key creation
severity: high
enforce: true
expression: |
event.tool.name == 'create_file' &&
event.tool.args.path.matches('.*/\\.ssh/id_.*') &&
!event.tool.args.path.endsWith('.pub')
message: Agent attempted to write SSH private key
The enforce: true flag determines whether findings trigger blocking or just logging. Operators can override shipped rules by placing custom YAML in --rules-dir with higher version numbers—the versioning prevents accidental downgrades when updating the binary. Output is versioned NDJSON with JSON Schema contracts, designed to stream into existing log infrastructure without custom collectors.
Gotcha
Coverage is the elephant in the room. As of the current release, only 4-5 agents support live pre-action hooks. Most entries in the compatibility matrix are 'forensic reconstruction only'—you can parse session logs after the fact, but you can't block actions in real-time. Major platforms like Cursor and Aider are marked 'deferred,' which is a polite way of saying 'we haven't built it yet.' If you're betting on Numbat to protect a diverse agent fleet, check the compatibility matrix first—you might find your primary agent only supports post-hoc analysis.
The enforcement model is fundamentally cooperative. When Numbat sends a deny decision, it trusts the agent to respect it. There's no OS-level backstop, no eBPF enforcing syscall policy, no kernel module. If an agent misbehaves (ignoring deny responses) or if an attacker modifies the agent runtime to skip hook calls entirely, Numbat sees nothing. The documentation is refreshingly honest about this: findings 'do not prove either action completed.' This isn't a cryptographic guarantee—it's detective controls masquerading as preventive. For truly hostile scenarios, you need defense in depth: Numbat catches mistakes and misconfigurations, not determined adversaries with code execution on the endpoint.
Verdict
Use Numbat if you're deploying AI coding agents in production environments with access to real credentials, need compliance evidence of agent activity, and already have infrastructure to consume NDJSON log streams. The pre-action blocking for high-confidence rules (SSH keys, cloud metadata, secrets patterns) provides a safety net that post-execution EDR can't match, and the canonical event model future-proofs detection logic as you add more agents. Skip it if your agents run in sandboxed dev containers where exfiltration is contained anyway, if you lack the eng resources to integrate NDJSON outputs with your SIEM, or if your threat model includes sophisticated attackers who'll bypass userland hooks. Also skip if you're still evaluating agents—the instrumentation overhead and trust-model friction aren't worth it until agents graduate from experiments to tools that touch production systems.