> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

Building Self-Improving AI Workflows with Claude Code's Hook System

[ View on GitHub ]

Building Self-Improving AI Workflows with Claude Code's Hook System

Hook

What if your AI coding assistant could automatically learn from its mistakes, remember context across sessions, and delegate specialized tasks to cheaper models—all without a single API call?

Context

Anyone who's used AI coding assistants for more than a week hits the same wall: you're constantly re-explaining your project structure, pasting the same coding standards, and watching your context window fill with irrelevant analysis. Claude Code closes and you lose everything. Start a new task and you're copy-pasting instructions from a text file like it's 2010.

The everything-claude-code toolkit attacks this friction by treating Claude Code not as a finished product but as a scriptable platform. Instead of fighting the 200k token limit, it implements surgical context management through event-driven hooks. Instead of manually maintaining instruction files, it extracts patterns from successful sessions into reusable skills. Instead of bloating your main context with security reviews and architectural analysis, it delegates to specialized subagents running on cheaper models. It's essentially a plugin system for Claude Code that adds memory, automation, and learning capabilities that should arguably be built-in.

Technical Insight

The architecture revolves around five primitives that interlock in clever ways. At the foundation are rules—markdown files that get injected into every system prompt. These establish baseline behavior like "Always run tests after code changes" or "Use TypeScript strict mode." Simple, but always active.

Hooks are where it gets interesting. These are Node.js scripts triggered by PreToolUse, PostToolUse, or Stop events, filtered through JSONLogic matchers. Here's a real example from the codebase that warns about console.log statements only in production code:

// hooks/warn-console-log.js
module.exports = {
  trigger: {
    event: "PreToolUse",
    condition: {
      "and": [
        {"==": [{"var": "tool.name"}, "str_replace_editor"]},
        {"in": ["console.log", {"var": "tool.params.new_str"}]},
        {"!": {"in": [".test.", {"var": "tool.params.path"}]}}
      ]
    }
  },
  execute: async (context) => {
    return {
      warning: "Detected console.log in production code. Consider using a proper logging library.",
      suggested_action: "Replace with logger.info() or remove before commit"
    };
  }
};

This fires before Claude applies the edit, injecting a warning into the conversation. The JSONLogic matcher is checking tool name, searching the replacement string for console.log, and excluding test files—all declaratively. No need to pollute Claude's main instructions with "don't forget to check for console.log" when you can intercept it surgically at edit time.

The agent delegation pattern solves context explosion. Instead of asking Claude to do security review in the same conversation where it's writing features, you invoke a specialized agent:

// agents/security-reviewer.json
{
  "name": "security-reviewer",
  "model": "claude-3-haiku-20240307",
  "tools": ["read_file", "list_files"],
  "skills": ["security-audit", "owasp-top-10"],
  "rules": ["security-focused"],
  "temperature": 0.2
}

This subagent runs on Haiku (cheaper, faster), can only read files (can't modify), and has constrained context loaded from security-specific skills. When you run /review-security, it spawns this agent, passes the relevant file paths, gets structured output, then returns a summary to your main session. Your primary Claude instance never sees the detailed OWASP analysis—it just gets "3 issues found, see security-report.md."

Memory persistence is implemented through Stop event hooks that serialize conversation state:

// hooks/persist-session.js
module.exports = {
  trigger: { event: "Stop" },
  execute: async (context) => {
    const sessionData = {
      id: context.sessionId,
      timestamp: Date.now(),
      messages: context.conversation.messages,
      metadata: {
        filesModified: context.session.filesModified,
        toolsUsed: context.session.toolsUsed,
        workingDirectory: context.workingDirectory
      }
    };
    
    await fs.writeFile(
      `${process.env.HOME}/.claude/sessions/${context.sessionId}.json`,
      JSON.stringify(sessionData, null, 2)
    );
  }
};

When you restart Claude, a PreSession hook loads this file and reconstructs context. It's not true persistence—it's replaying the conversation history—but it means long-running projects don't lose critical context every time you close the editor.

The continuous learning loop is perhaps most ambitious. After completing a session, a hook analyzes the transcript:

// Simplified from evaluate-session.js
async function extractPatterns(session) {
  const patterns = [];
  
  // Find tool usage sequences that succeeded
  const successfulFlows = session.messages
    .filter(m => m.toolUse && m.followingTest?.passed)
    .map(m => ({
      context: m.precedingMessages.slice(-3),
      tools: m.toolSequence,
      outcome: m.result
    }));
  
  // Convert to skill definitions
  for (const flow of successfulFlows) {
    if (isNovel(flow) && hasGeneralizablePattern(flow)) {
      patterns.push({
        skill: generateSkillMarkdown(flow),
        confidence: calculateConfidence(flow)
      });
    }
  }
  
  return patterns;
}

This extracts sequences like "when fixing TypeScript errors, run type check before and after changes" into skill files that future sessions can reference. Over time, your Claude instance builds a library of patterns specific to your codebase.

The package manager detection showcases practical cross-platform engineering. Instead of hardcoding npm install, it cascades through multiple heuristics:

function detectPackageManager() {
  // 1. Explicit env var
  if (process.env.CLAUDE_PKG_MANAGER) return process.env.CLAUDE_PKG_MANAGER;
  
  // 2. Project config
  const packageJson = readPackageJson();
  if (packageJson?.packageManager) {
    return packageJson.packageManager.split('@')[0]; // "pnpm@8.0.0" -> "pnpm"
  }
  
  // 3. Lockfile sniffing
  if (fs.existsSync('pnpm-lock.yaml')) return 'pnpm';
  if (fs.existsSync('yarn.lock')) return 'yarn';
  if (fs.existsSync('package-lock.json')) return 'npm';
  
  // 4. Global config
  const globalConfig = execSync('npm config get package-manager', { encoding: 'utf8' });
  if (globalConfig && globalConfig !== 'undefined') return globalConfig.trim();
  
  // 5. First available
  for (const pm of ['pnpm', 'yarn', 'npm']) {
    if (commandExists(pm)) return pm;
  }
  
  throw new Error('No package manager detected');
}

This handles monorepos, alternative package managers, and containerized environments where global installs might differ from project expectations.

Gotcha

The hook system's synchronous execution model creates real latency problems. Every file edit triggers PreToolUse hooks, which block until completion. If your hook runs grep -r "TODO" . across a large codebase, Claude freezes until it completes. There's no timeout mechanism, no async execution, no way to parallelize multiple hooks. You'll feel this immediately in repositories with thousands of files.

Memory persistence is fragile. The session files are plaintext JSON in your home directory with no encryption, meaning sensitive data from conversations (API keys discussed, architecture decisions, database schemas) sits readable by any process. Concurrent sessions clobber each other's state files since there's no file locking. If Claude crashes mid-session, you lose everything since the Stop hook never fires. The rehydration process replays the entire conversation history, which counts against your token budget—sessions with 500 messages mean you're paying to re-process all that context on startup.

The continuous learning pattern has no quality control. Failed sessions get mined for patterns just as eagerly as successful ones. There's no deduplication, so "run tests after changes" gets extracted into five different skill files with slight variations. The confidence scoring is simplistic—it counts tool usage frequency but doesn't actually validate that the pattern generalizes. After a few weeks, you'll have dozens of skill files with overlapping advice and no clear way to prune them.

Finally, this only works with Claude Code. The entire plugin manifest system, the hook triggers, the agent delegation—it's all specific to Anthropic's editor. If you're using Cursor, Continue.dev, or raw API access, none of this applies. You're locked into their platform and their update cadence.

Verdict

Use if: You're doing serious development work in Claude Code for projects spanning multiple sessions, you're already manually copy-pasting instructions or fighting context window limits, or you're on a team that needs reproducible AI workflows with standardized skills and rules. The hook system and agent delegation genuinely solve problems you'll hit after a few weeks of heavy use, and the memory persistence is invaluable for long-running projects. Also useful if you want to experiment with meta-learning patterns where AI tools improve their own configuration over time.

Skip if: You're platform-agnostic and might switch to Cursor or other editors (this is hyper-specific to Claude Code), you prefer explicitly controlling every AI interaction rather than automated hooks firing in the background, you're just experimenting with AI coding tools on small projects where context resets aren't painful, or you don't want the maintenance burden of curating skill files and debugging hook execution. The added complexity only pays off if you're committed to Claude Code long-term and willing to treat your configuration as living infrastructure that needs occasional pruning and refinement.