> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

bb: The Agent IDE That Treats Workflows as First-Class Resumable Primitives

[ View on GitHub ]

bb: The Agent IDE That Treats Workflows as First-Class Resumable Primitives

Hook

Most AI coding tools trap you in a single-agent conversation that evaporates when you close the window. bb treats agent workflows as persistent, resumable threads that survive handoffs between humans, Claude, and custom automation—a fundamentally different architecture that most developers haven't encountered yet.

Context

The first wave of AI coding assistants—GitHub Copilot, Cursor, Windsurf—solved the single-agent problem beautifully. You chat with Claude or GPT-4, it edits your code, you accept or reject changes. This works remarkably well until you hit the boundaries of the paradigm: What happens when you need to pause an agent mid-task and hand off to a human reviewer? Can you orchestrate multiple agents across different tools working on the same codebase? Can you programmatically trigger agent workflows from CI/CD or internal tooling?

These questions reveal a deeper architectural constraint: most AI IDEs are monolithic applications with opaque execution models. The agent lives inside the UI, state exists only during the session, and programmatic access is an afterthought if it exists at all. bb emerged from a different premise entirely—what if the agent runtime was the product, and the IDE was just one client among many? What if workflows persisted as first-class database entities that any surface could read, write, and resume? This API-first, multi-surface architecture enables use cases the first generation of tools structurally cannot support: CI bots that start agent tasks, CLIs that query thread status, web dashboards that visualize cross-project agentic work, and most provocatively, agents that modify bb's own codebase through bb's own API—a meta-circular development loop that justifies the tagline 'the agent IDE that builds itself.'

Technical Insight

Workspace

Core

Surfaces

Commands

Commands

Commands

Manage State

Persist

Read State

WebSocket Push

WebSocket Push

WebSocket Push

Watch Changes

Coordinate

Execute Tasks

Desktop App

Electron

Web UI

Vite

CLI Tool

HTTP/WebSocket API

Node.js Server

Thread Manager

SQLite DB

Threads/Tasks/Messages

Host Daemon

Filesystem Watcher

Project Files

System architecture — auto-generated

bb's architecture centers on a long-running Node.js server backed by SQLite that treats threads, tasks, and messages as durable state. Unlike Cursor where conversations live in memory and vanish on restart, bb persists every interaction. A thread represents a complete agentic workflow—code generation, file modifications, test runs—that can be paused, inspected, and resumed across process boundaries. This isn't just logging; it's transactional state management for agent execution.

The codebase is a TypeScript monorepo with three core components: the server process exposes HTTP and WebSocket APIs for thread management, the host daemon watches filesystem changes and coordinates project workspaces, and multiple surfaces (Electron desktop app, Vite web UI, CLI) consume the same API endpoints identically. Communication flows unidirectionally: surfaces send commands via HTTP, the server updates SQLite state, and WebSocket pushes broadcast changes back to connected clients. This clean separation means adding a new surface—say, a Slack bot or VS Code extension—requires only implementing the API client, not touching core execution logic.

Here's what creating and resuming a thread looks like via the HTTP API:

// Start a new thread with an agent task
const response = await fetch('http://localhost:33900/api/threads', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    projectId: 'my-project',
    messages: [{
      role: 'user',
      content: 'Refactor the authentication module to use JWT'
    }]
  })
});

const { threadId } = await response.json();

// Later, from a different process or even different machine:
// Query thread status
const status = await fetch(`http://localhost:33900/api/threads/${threadId}`);
const thread = await status.json();

// Resume execution by appending a message
await fetch(`http://localhost:33900/api/threads/${threadId}/messages`, {
  method: 'POST',
  body: JSON.stringify({
    role: 'user',
    content: 'Actually, use refresh tokens too'
  })
});

This API-first design inverts the typical IDE architecture. In Cursor, the UI is the product and the agent lives inside it. In bb, the server managing thread state is the product, and the desktop app is just one client. This matters enormously for orchestration: you can script agent workflows, chain tasks across multiple agents, or build dashboards that visualize agentic work across an entire organization—all impossible when the execution model is trapped inside a UI process.

The deterministic dev environment setup reveals another architectural decision that solves a problem most tools ignore. When you run multiple git worktrees or checkout the same project in different directories, you need isolated dev servers that don't conflict. bb hashes the checkout path and derives port numbers deterministically—the same directory always gets the same ports. The server runs on 33900 + hash(path) % 1000, the Vite dev server offsets differently, and the host daemon gets its own range. This means you can run three different bb branches simultaneously, each with its own isolated database and server, without manual port configuration. It's the kind of pragmatic detail that only emerges from real multi-worktree development workflows.

The deliberate rejection of hot-reload for the server and daemon components is equally revealing. The Vite app hot-reloads, but server code changes require explicit restarts. This isn't laziness—it's architectural honesty. When you're managing stateful agent threads mid-execution, hot-reloading could corrupt in-flight workflows. An agent half-done with a refactoring task represents transactional state that can't safely reload. The team chose development friction over runtime fragility, prioritizing correctness for long-running agentic processes over developer convenience. This trade-off is invisible in stateless tools but critical in bb's durable-workflow model.

Native addon dependencies on better-sqlite3 and @parcel/watcher signal performance requirements that pure-JavaScript solutions couldn't meet. Watching large codebases for changes needs native filesystem APIs, and maintaining transactional integrity across concurrent agent operations hitting the database demands SQLite's locking guarantees through native bindings. The cost is deployment friction—any environment blocking native compilation breaks bb entirely—but the team chose performance for core use cases over universal compatibility.

Gotcha

The macOS Apple Silicon exclusivity for desktop builds exposes resource constraints. Windows requires WSL2, Linux users run via npx, and Intel Mac support is uncertain. This isn't a polished cross-platform product; it's a tool built by a small team scratching their own itch on M1 hardware. If your team is on varied platforms, expect friction or stick to the web UI and CLI.

SQLite as the sole persistence layer creates hard scaling limits. This works beautifully for single developers or small co-located teams sharing a filesystem, but there's no clustering, no replication, no multi-tenant story. If you're imagining bb as infrastructure for a 50-person engineering org, you'll hit walls immediately. The database is a file on disk—coordination beyond that requires manual syncing or shared network drives, neither of which is robust. The architecture fundamentally doesn't scale beyond laptop-class deployments, which disqualifies most enterprise scenarios.

The 'actively evolving' disclaimer in the README is both honest and concerning. Core thread primitives are stable, but workflows and surfaces are still churning. Early adopters should expect breaking changes, incomplete documentation, and features that work differently month-to-month. This is a tool for teams comfortable reading source code and adapting to churn, not for organizations needing API stability guarantees.

Verdict

Use if: You're building agent orchestration infrastructure and need programmatic control over AI coding workflows—the HTTP API and resumable threads enable automation patterns impossible in Cursor or Windsurf. Use if you're experimenting with multi-agent systems where different tools need to hand off work without losing context. Use if you're a TypeScript shop comfortable running your own services and troubleshooting native addon compilation issues. Skip if: You want a polished, single-agent coding experience—Cursor is more mature and Windsurf has better UX for that use case. Skip if your team uses Windows natively or mixed platforms—the macOS-first build strategy will cause constant friction. Skip if you need multi-user coordination beyond shared filesystems—SQLite's limitations make this a non-starter for team-scale deployments. Skip if you need production-grade stability—this is actively evolving infrastructure for early adopters who can tolerate churn.