> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

Agent Swarm: The Multi-Agent Framework That Actually Remembers What It Learned

[ View on GitHub ]

Agent Swarm: The Multi-Agent Framework That Actually Remembers What It Learned

Hook

Most AI agent frameworks treat every task like it's their first day on the job. Agent Swarm builds a framework where agents genuinely get smarter over time—not through bigger context windows, but through a compounding memory system that survives deployments and creates institutional knowledge.

Context

If you've deployed LangChain agents in production, you've hit the wall: your customer support bot answers the same question differently every time. Your code review agent ignores the style guide it 'learned' yesterday. Your workflow automation forgets how it solved an edge case last week. The problem isn't hallucination—it's amnesia.

Agent frameworks like AutoGen and CrewAI excel at one-shot tasks but treat memory as an afterthought. They stuff everything into context windows or dump conversation history into vector databases without structure. When the agent restarts, that knowledge evaporates. When you scale to multiple agents, there's no shared institutional brain. Desplega AI built Agent Swarm to solve this: a TypeScript-based agentic operating system where memory compounds across tasks, agents inherit learnings from their predecessors, and complex multi-stage workflows can span days without losing state.

Technical Insight

Agent Swarm's architecture centers on three interlocking systems: hierarchical orchestration, harness abstraction, and what they call compounding memory.

The orchestration layer implements a lead-worker pattern. A lead agent receives tasks through ingestion channels—Slack webhooks, GitHub events, Linear tickets, even IMAP email polling. It decomposes complex requests into subtasks and spawns worker agents in isolated Docker containers. Each worker executes via a pluggable harness: Anthropic's Claude Code, OpenAI Codex, a custom Devin integration, or raw LLM calls routed through pi-mono (which abstracts Anthropic, OpenRouter, and AWS Bedrock). This harness abstraction is clever—you configure tasks with intent tiers like 'smol', 'regular', 'smart', or 'ultra', and the system maps these to whatever models you've provisioned. When GPT-4 gets deprecated, you update one config file instead of rewriting task definitions.

Here's how you'd define a workflow that spawns parallel workers:

import { defineWorkflow } from '@agent-swarm/sdk';

export default defineWorkflow({
  name: 'code-review-pipeline',
  steps: [
    {
      id: 'analyze-pr',
      type: 'agent-task',
      harness: 'claude-code',
      modelTier: 'smart',
      prompt: 'Review PR {{inputs.prUrl}} for security issues'
    },
    {
      id: 'parallel-checks',
      type: 'foreach',
      items: '{{steps.analyze-pr.output.files}}',
      workflow: {
        steps: [
          {
            id: 'lint-file',
            type: 'agent-task',
            harness: 'openai-codex',
            modelTier: 'smol',
            prompt: 'Run linter on {{item.path}}'
          }
        ]
      }
    },
    {
      id: 'approval-gate',
      type: 'hitl',
      prompt: 'Approve deployment? Issues: {{steps.parallel-checks.summary}}'
    }
  ]
});

This workflow spawns N worker containers for parallel linting, pauses at a human-in-the-loop approval gate, and resumes when you click a Slack button. The DAG engine tracks dependencies—parent steps wait for all child tasks to complete before proceeding. State persists in PostgreSQL, so if your cluster restarts mid-workflow, tasks resume from the last checkpoint.

The compounding memory system is where Agent Swarm diverges from typical RAG implementations. Instead of dumping text chunks into pgvector and hoping semantic search works, it maintains structured memory types: agent identity files (SOUL files, persona updates), linked memory graphs, task outcomes with usefulness scores, and version-controlled amendments. When an agent recalls memories, it runs a hybrid query combining pgvector similarity search with PostgreSQL full-text ranking, then expands results by following graph links to related memories.

The critical insight: agents don't just read from this memory corpus—they actively curate it. After completing a task, an agent can tag its learnings with usefulness scores, link new memories to existing ones ("this solution extends the pattern we used in task #423"), or amend previous memories without deleting them (preserving provenance). Over weeks, this creates an institutional knowledge base. Your support agent doesn't just answer "how do I reset my password"—it recalls that you updated the password reset flow three times, references the latest design doc, and notes that Enterprise customers have SSO exceptions.

The filesystem layer (agent-fs) handles artifacts too large for database BLOBs—think multi-gigabyte model checkpoints or video files. It exposes a POSIX-like interface to object storage (S3/GCS), with namespacing that inherits from task contexts. An agent processing customer videos in task #500 can only see files in that task's namespace unless explicitly granted cross-namespace access.

Extensibility happens through the Model Context Protocol (MCP). Agents get tools like memory.recall, kv.get, workflow.trigger, and script.run. That last one is meta-programmable: agents write TypeScript scripts that get JIT-compiled and executed with full SDK access. A script can query memories, call external APIs, transform data, then publish itself as a versioned HTTP endpoint with bearer auth. Effectively, agents build their own tools:

// Agent-written script that becomes a tool
export default async function analyzeChurn(ctx) {
  const recentTickets = await ctx.memory.recall({
    query: 'customer cancellation reason',
    limit: 50,
    since: '30d'
  });
  
  const patterns = await ctx.llm.analyze({
    prompt: 'Cluster these cancellation reasons',
    data: recentTickets
  });
  
  await ctx.memory.store({
    type: 'insight',
    content: patterns,
    tags: ['churn-analysis'],
    linkedTo: recentTickets.map(t => t.id)
  });
  
  return patterns;
}

Once published, other agents can invoke this as custom.analyzeChurn(), and the memory linkage means future recalls surface this analysis alongside raw ticket data.

The dashboard deserves mention. Instead of polling APIs, it subscribes to PostgreSQL LISTEN/NOTIFY channels via WebSockets. When an agent updates task status or stores a memory, the database fires events that stream to connected clients in real-time. The UI renders live task trees—you watch agents spawn children, fork parallel branches, hit approval gates, and rejoin. For Slack integrations, Agent Swarm avoids notification spam by editing a single outcome card per conversation thread rather than posting every tool invocation, while maintaining a hidden task tree structure for debugging.

Gotcha

The Docker-per-task execution model creates brutal cold-start latency. Spinning up a worker container, pulling images, and initializing the harness can take 5-15 seconds. If you're building a conversational chatbot where users expect sub-second responses, this is a non-starter. Serverless frameworks like LangGraph Cloud get you sub-100ms cold starts by keeping runtimes warm. Agent Swarm optimizes for correctness and isolation over speed—it's built for workflows that take minutes or hours (generate report, review 20 PRs, process support backlog), not interactive chat.

PostgreSQL as the single source of truth will bottleneck at scale. Memory recall queries hit pgvector sequentially. Workflow state updates require transactions. The LISTEN/NOTIFY pattern for real-time dashboards breaks if you horizontally shard your database. The system is documented to handle "dozens of concurrent tasks," not thousands. If you need to scale beyond a single beefy Postgres instance, you're rewriting core persistence logic.

The compounding memory system lacks garbage collection. Agents accumulate learnings indefinitely with no documented pruning strategy. After six months of operation, your memory corpus might contain 100,000 entries. Does vector recall still surface relevant results, or does it drown in noise? The repo doesn't address relevance decay, archival policies, or how to migrate old memories when you refactor your agent personas. You'll need to implement this operational discipline yourself, or watch retrieval quality degrade over time.

Verdict

Use Agent Swarm if you're building internal AI tooling for platform engineering teams where institutional knowledge compounds—automating Linear→PR workflows that learn coding standards, support agents that accumulate product knowledge across quarters, or compliance bots that evolve with regulation changes. The memory system and workflow orchestration justify the operational complexity when tasks span days and context survival matters more than response latency. You need devops capacity to run Postgres, Redis, Docker, and object storage, plus appetite to debug container networking when things break. Skip this entirely if you need sub-second response times (use LangGraph Cloud), want serverless deployment (try CrewAI or AutoGen), operate in regulated environments where agent-written code execution raises compliance red flags, or lack infrastructure experience—this is a database-backed orchestration platform that happens to run AI agents, not a lightweight framework you npm install and forget.