> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

Eve: The Agent Framework That Thinks Your Filesystem Is an API

[ View on GitHub ]

Eve: The Agent Framework That Thinks Your Filesystem Is an API

Hook

Most agent frameworks make you learn a new SDK to add capabilities. Eve lets you drop a TypeScript file in a folder and restart the server.

Context

Building production agents in 2024 means choosing between Python frameworks drowning in abstractions (LangChain's 47 different chain types) or rolling your own plumbing around streaming responses, function calling, and conversation state. The first-generation frameworks treated agents as research projects—you'd wire together retrievers, memory modules, and planning loops with enough YAML to make a Kubernetes engineer nostalgic. The result was opaque systems where adding a single tool required touching five files and understanding the framework's entire mental model.

Eve takes a radically different approach: your agent's capabilities are defined by what files exist in specific directories. Want to add a tool? Create tools/weather.ts. Need to give the agent domain knowledge? Drop skills/customer-data.md in the skills folder. This convention-over-configuration philosophy borrows from Next.js's pages router—the filesystem structure is the configuration. For teams already on Vercel, it eliminates the impedance mismatch between their frontend conventions and their agent backend. For everyone else, it's a bet that discoverability (you can understand the agent by running ls) matters more than programmatic flexibility.

Technical Insight

Eve's architecture revolves around a directory scanner that hot-loads modules at startup. The runtime watches four key directories: tools/ for synchronous functions, skills/ for Markdown procedures, channels/ for I/O adapters, and schedules/ for cron triggers. Each has a specific TypeScript interface that the framework expects. Here's what a tool definition looks like:

import { z } from 'zod';
import { tool } from 'eve';

export default tool({
  name: 'search_documentation',
  description: 'Searches internal docs for the given query',
  parameters: z.object({
    query: z.string().describe('The search term'),
    limit: z.number().optional().default(5)
  }),
  execute: async ({ query, limit }) => {
    const results = await docSearch.query(query, limit);
    return { results };
  }
});

The Zod schema does double duty: it generates TypeScript types for your execute function and runtime validation for the LLM's tool calls. When the model hallucinates a parameter or gets a type wrong, Eve catches it before execution and returns a validation error to the model, letting it retry. This is architecturally superior to Python frameworks where type hints are decorative and you're debugging stringified JSON dumps in production.

The skills vs. tools distinction solves a real context window economics problem. Tools are always injected into the system prompt—they're cheap, synchronous functions like "get current time" or "validate email format." Skills are heavyweight: multi-paragraph procedures, domain knowledge, or complex workflows stored as Markdown. They're only loaded when relevant, presumably via semantic search over embeddings or keyword matching (Eve's docs are maddeningly vague here). A skill looks like this:

# Refund Processing

When a customer requests a refund:

1. Verify the order ID exists in Stripe
2. Check refund eligibility (must be <30 days)
3. If eligible, call the `process_refund` tool
4. Send confirmation email via `send_email` tool
5. Log the interaction to the CRM

Do not process refunds >$500 without manager approval.

The agent runtime decides when to load this into context. If a user says "I want a refund," the system retrieves relevant skills, injects them into the conversation, and the LLM follows the procedure. This is elegant in theory—you're externalizing business logic from code into prose that non-engineers can edit. In practice, it's another AI-powered retrieval system with all the usual failure modes: irrelevant skill injection, missed edge cases, and zero guarantees about deterministic execution.

Channels are where Eve's production chops show. Instead of building an HTTP endpoint and bolting on Slack later, you define channel adapters that normalize different surfaces into Eve's internal message format:

import { channel } from 'eve';
import { WebClient } from '@slack/web-api';

export default channel({
  name: 'slack',
  setup: async () => {
    return new WebClient(process.env.SLACK_TOKEN);
  },
  handleMessage: async (client, event) => {
    return {
      conversationId: event.channel,
      userId: event.user,
      text: event.text,
      respond: async (message) => {
        await client.chat.postMessage({
          channel: event.channel,
          text: message
        });
      }
    };
  }
});

This abstraction means your agent logic is surface-agnostic. The same tools and skills work via HTTP, Slack, Discord, or a CLI—you're not littering your code with if (platform === 'slack') conditionals. The framework handles streaming responses, typing indicators, and message formatting per platform.

Under the hood, Eve wraps Vercel's AI SDK, which provides model abstraction, streaming via Server-Sent Events, and automatic retries with fallback providers. The SDK call looks like this internally:

const response = await streamText({
  model: openai('gpt-4-turbo'),
  messages: conversationHistory,
  tools: loadedTools,
  maxSteps: 5
});

The maxSteps parameter is doing heavy lifting—it's the agent loop. The SDK automatically handles the model calling a tool, executing it, appending results to history, and calling the model again until it produces a text response instead of a tool call. This is the orchestration Eve doesn't provide: it's delegated to the AI SDK's built-in loop. If you need complex branching ("retry tool X three times with exponential backoff, then call tool Y, then human-in-the-loop approval"), you're implementing that logic inside tool functions, not as first-class workflow primitives.

Gotcha

The filesystem-as-API paradigm breaks down when you need programmatic agent composition. You can't easily spin up multiple agents with different tool sets in the same process or dynamically generate agent configurations based on user permissions. The framework assumes one agent per deployment, which is fine for "customer support bot" but limiting for "marketplace where each seller has a custom agent" architectures. You'd need to run multiple Eve instances or abandon the filesystem conventions entirely and use the AI SDK directly.

The Vercel coupling is the elephant in the deployment story. Eve depends on Vercel AI Gateway for model routing, which means every LLM call goes through Vercel's infrastructure with their markup. The README claims it's "open," but there's no escape hatch to swap in direct OpenAI/Anthropic API calls without rewriting channel and runtime code. If you're already paying for Vercel hosting, this is synergy. If you're on AWS or running on-prem, it's a dealbreaker. The durability claims are also unsubstantiated—there's no documentation on how conversation state persists, what happens on crashes, or how to migrate between versions. For a framework targeting production use, these omissions are glaring.

Verdict

Use if: You're shipping an agent on Vercel infrastructure and value velocity over flexibility—the filesystem conventions eliminate boilerplate, the AI SDK handles streaming/retries, and channels let you serve multiple surfaces from one codebase. Also choose this if you need non-engineers to edit agent behavior (Markdown skills are more approachable than Python code) or want interpretability (the agent's capabilities are literally ls tools/).

Skip if: You need vendor-neutral deployment, fine-grained control over LLM APIs, or complex orchestration beyond what models can reason about. The Vercel lock-in isn't just inconvenient—it's architectural. Also skip if you're building multi-agent systems or need programmatic agent composition; the filesystem conventions assume one agent per deployment and don't support dynamic capability loading.