OpenExecutive: How Prompt Caching and Episodic Memory Turn 8 Claude Agents Into a Persistent Strategic Advisor
Hook
Most multi-agent systems forget everything between sessions. OpenExecutive runs a background Haiku pass after every conversation to mine structured decisions from unstructured dialogue—building a memory layer that surfaces last quarter's strategic recommendations without you asking.
Context
If you've tried building an AI advisor for strategic work, you've hit the same wall: ChatGPT is stateless, custom agents are expensive to run at scale, and most frameworks treat every conversation as a cold start. You end up copy-pasting context from previous threads or manually briefing the AI on what it already 'knows.' Enterprise teams resort to Notion pages titled 'What We Told the AI Last Time' because no system remembers on its own.
OpenExecutive attacks this with two architectural bets. First, it splits a single executive persona across eight domain specialists (CSO, CFO, CHRO, General Counsel, COO, CMO, CPO, Board Advisor) that retrieve from a git-tracked MBA knowledge base and company-uploaded documents. Second, it extracts episodic memory—decisions made, initiatives started, follow-ups pending—into SQLite after every response, then injects that context into future sessions. The result is an AI that behaves less like a chatbot and more like an executive assistant with institutional memory.
Technical Insight
The orchestration layer is deceptively simple: user messages hit an Executive Orchestrator running claude-sonnet-4-6, which uses parallel function calls to route questions to specialist agents. Each specialist is a separate Claude API call with domain-tuned system prompts and retrieval from two ChromaDB collections—one for builtin business knowledge (markdown files on financial modeling, GTM strategy, employment law), one for company-specific documents (pitch decks, board decks, financial models).
The non-obvious part is how RAG context gets injected. Instead of stuffing retrieval results into the system prompt—which would break Anthropic's prompt caching on every query—OpenExecutive injects chunks into the user turn:
# Retrieve relevant context from ChromaDB
builtin_results = builtin_collection.query(
query_texts=[user_message],
n_results=5
)
company_results = company_collection.query(
query_texts=[user_message],
n_results=3
)
# Inject into user message, preserving static system prompt
augmented_user_message = f"""
<retrieved_context>
{format_chunks(builtin_results)}
{format_chunks(company_results)}
</retrieved_context>
<user_query>
{user_message}
</user_query>
"""
# System prompt stays unchanged, enabling prompt caching
response = anthropic.messages.create(
model="claude-sonnet-4-6",
system=static_specialist_prompt, # Cached across requests
messages=[{"role": "user", "content": augmented_user_message}]
)
This tradeoff sacrifices retrieval visibility—the model can't see what was retrieved when forming its 'reasoning trace'—but achieves 85% cache hit rates after warmup because the system prompt (persona definition, company profile, knowledge index structure) never changes. For a deployment handling dozens of questions daily, that's the difference between $2 and $30 in API costs per conversation.
Episodic memory extraction runs as a background task after every specialist response. A separate Haiku call parses the conversation turn and outputs structured JSON:
async def extract_episodic_memory(conversation_turn: dict):
extraction_prompt = """
Extract decisions, commitments, and action items from this exchange.
Output JSON with: decision_made, rationale, stakeholders, follow_up_date.
"""
memory_response = await anthropic.messages.create(
model="claude-haiku-3-5",
messages=[{
"role": "user",
"content": f"{extraction_prompt}\n\n{conversation_turn}"
}]
)
# Parse and store in SQLite
memory_entries = json.loads(memory_response.content)
for entry in memory_entries:
db.execute("""
INSERT INTO episodic_memory
(decision, rationale, stakeholders, follow_up_date, session_id)
VALUES (?, ?, ?, ?, ?)
""", (entry['decision_made'], entry['rationale'],
json.dumps(entry['stakeholders']),
entry['follow_up_date'], session_id))
Future sessions query this table and inject a <past_decisions> block into the system prompt during onboarding. The model sees 'Three months ago, the CSO recommended focusing on enterprise over SMB; CFO flagged 18-month runway as the constraint' without the user re-explaining context. The memory layer emerges from conversational analysis rather than explicit commands like 'remember this.'
The scheduler uses SQLite's UPDATE...RETURNING to atomically claim jobs:
# Claim next due action without race conditions
claimed_task = db.execute("""
UPDATE scheduled_actions
SET status = 'claimed', claimed_at = ?
WHERE id = (
SELECT id FROM scheduled_actions
WHERE status = 'pending' AND due_at <= ?
ORDER BY due_at ASC
LIMIT 1
)
RETURNING *
""", (datetime.now(), datetime.now())).fetchone()
This Postgres-style atomicity works on SQLite 3.35+ and prevents double-execution without distributed locks. The tradeoff is hard: you cannot run multiple backend instances. Horizontal scaling requires ripping out SQLite for Postgres and adding a proper job queue like Celery or Temporal.
Integration adapters (Slack, Discord, email IMAP poller) run as FastAPI lifespan context managers—they start when the server boots and share the same database connections. A Discord message hits the same orchestration pipeline as a CLI query, and the episodic memory extracted from a Slack thread informs answers delivered via email. This architectural choice collapses failure domains (Discord bot crash kills the API server) but eliminates the operational overhead of deploying separate services for each channel.
Gotcha
The single-instance architecture is a hard ceiling. You cannot horizontally scale OpenExecutive without replacing SQLite with Postgres, externalizing ChromaDB to a standalone server (losing embedding model consistency guarantees), and refactoring the scheduler to use a distributed job queue. The README doesn't surface this constraint clearly—if you deploy this for a portfolio of companies or expect sustained traffic beyond one executive team's usage, you'll hit concurrency limits and have no migration path except a rewrite.
Episodic memory extraction trusts Haiku to accurately parse unstructured conversation into structured facts. There's no validation that extracted decisions are correct, no conflict resolution when the CSO's recommendation contradicts the CFO's, and no semantic deduplication. After a dozen sessions, you may see 'Focus on enterprise' recorded three times with slightly different wording, and the system has no mechanism to reconcile or canonicalize. The memory layer is append-only with no garbage collection—long-running deployments will inject increasingly noisy context into the <past_decisions> block.
RAG quality is bottlenecked by sentence-transformers and cosine similarity. There's no hybrid search combining semantic and keyword matching, no reranking, no query expansion. If your company docs use internal jargon that doesn't appear in the builtin MBA knowledge base, retrieval will miss relevant chunks. The builtin knowledge base is static markdown checked into Git—there's no versioning strategy for updating domain expertise without redeploying, and no way to A/B test knowledge changes across sessions.
Verdict
Use OpenExecutive if you're running a single company (startup, portfolio holding, or investment thesis tracker) and want a persistent strategic advisor accessible across Slack, email, and Discord. The episodic memory and proactive scheduler give it genuine workflow continuity that stateless ChatGPT wrappers lack—it remembers what it recommended last quarter and surfaces follow-ups without prompting. The builtin MBA knowledge base provides better structured reasoning than raw Claude API access, and the prompt caching architecture makes sustained usage economically viable. Skip it if you need multi-tenancy, horizontal scaling, or enterprise-grade reliability guarantees. The single-instance scheduler and volume-bound state make it architecturally unsuitable for SaaS products or deployments spanning multiple legal entities. Also skip if you need granular access control, audit logs, or SOC 2 compliance—the architecture assumes full trust and no isolation between users. Choose LangGraph or CrewAI instead if you need cyclic agent workflows, human-in-the-loop approvals, or the flexibility to swap orchestration patterns without forking the codebase.