> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

Token Envy: Privacy-First Performance Analytics for Claude Code Without Shipping Your Prompts

[ View on GitHub ]

Token Envy: Privacy-First Performance Analytics for Claude Code Without Shipping Your Prompts

Hook

Your Claude API responses are getting slower, but you can't prove it without sending your entire prompt history to a third-party analytics service. Token Envy solves this with a locally-encrypted SQLite database that tracks performance without ever exposing what you asked.

Context

If you're doing serious work with Claude Code—security research, extended prompt engineering sessions, or high-volume code generation—you've probably noticed response times varying wildly. Sometimes you get instant replies, other times you're waiting 30 seconds for the same request complexity. Is Anthropic throttling your account? Did you hit a quota ceiling? Or did your prompts just get more expensive?

The problem is proving it. Anthropic's web console shows account-level quota but nothing granular. Generic LLM observability tools like LangSmith or Helicone require routing all your traffic through their servers, which means shipping client work to third parties. When you're doing penetration testing or handling proprietary codebases, that's a non-starter. You need analytics that stay on your machine, understand Claude's specific quota systems and model variants, and can quantify performance shifts without storing raw prompts. Token Envy is the first tool purpose-built for this gap: a local-first analytics pipeline that watches Claude Code's JSONL transcript logs, computes mix-adjusted performance metrics, and pseudonymizes everything so aggressively that even a compromised database can't reverse-map to your source content.

Technical Insight

Token Envy's architecture is a four-layer pipeline that turns Claude Code transcripts into actionable performance metrics without ever exposing sensitive data. At the bottom, a filesystem watcher monitors ~/.claude/ for JSONL log files, parsing each request/response pair into structured records. These feed an SQLite aggregation engine that computes statistical summaries while pseudonymizing conversation IDs through HMAC-SHA256 digests keyed to your machine. A Node.js HTTP server exposes these aggregates via a token-authenticated API, and a browser frontend renders D3-style time-series visualizations.

The privacy model is architecturally enforced through irreversible pseudonymization. When Token Envy ingests a transcript line, it immediately discards the raw prompt and response text, keeping only metadata like token counts, model names, and timing:

const conversationDigest = createHmac('sha256', localMachineKey)
  .update(conversationId)
  .digest('hex')
  .substring(0, 16);

const record = {
  digest: conversationDigest,
  model: request.model,
  inputTokens: usage.input_tokens,
  outputTokens: usage.output_tokens,
  queueTime: timing.queue_ms,
  processingTime: timing.process_ms,
  generationTime: timing.generate_ms,
  timestamp: request.timestamp
};

This HMAC approach is critical. Unlike encryption, which can be reversed with a key, HMAC digests are one-way. Even if someone steals your SQLite database, they can't reconstruct conversation content or link metrics back to specific transcripts. The tool intentionally throws away the only data that could enable reverse lookup.

The standout technical feature is the mix-adjusted Speed Index, which normalizes performance across different Claude models. If you're alternating between Claude 3.5 Sonnet and Opus, raw timing averages are meaningless—Opus is architecturally slower but more capable. Token Envy solves this by computing a baseline distribution for each model (requiring 100+ requests) and then expressing each request as a percentile within that model's historical performance. The 28-day Speed Index averages these percentiles, weighted by request volume:

function computeMixAdjustedSpeed(requests: Request[]): number | null {
  const modelBaselines = getModelBaselines(requests);
  
  // Require 100+ requests and 70% coverage
  if (totalRequests < 100 || coverageRatio < 0.7) {
    return null;
  }

  let weightedSum = 0;
  let totalWeight = 0;

  for (const req of requests) {
    const baseline = modelBaselines[req.model];
    const percentile = computePercentile(req.effectiveSpeed, baseline);
    const weight = req.outputTokens;
    
    weightedSum += percentile * weight;
    totalWeight += weight;
  }

  return weightedSum / totalWeight;
}

This is sophisticated time-series analysis far beyond naive averaging. By tracking percentiles instead of absolute timings, you can detect when Anthropic's infrastructure degrades even as your own usage patterns shift between models. A Speed Index dropping from 50th to 30th percentile means you're getting slower responses relative to your own baseline, regardless of model mix.

The other clever piece is "effective output speed," which measures wall-clock experience rather than theoretical decoder throughput. It combines three phases: queueing time (waiting for capacity), processing time (prompt ingestion), and generation time (token production). Most benchmarks only measure generation, but real users care about end-to-end latency. If Anthropic's load balancers are congested, your queueing time spikes even though generation speed stays constant. Token Envy surfaces this in the UI as tokens per second from request submission to final token.

The status-line integration demonstrates a clever IPC pattern. Claude Code supports a statusLine setting that spawns a subprocess before each request. Token Envy installs itself here:

{
  "statusLine": "/usr/local/bin/tokenenvy-status --port 3000 --secret abc123"
}

When Claude Code invokes this, Token Envy parses the rate-limit headers from Claude's response and POSTs them to its own localhost server with the per-launch bearer token. This creates a side-channel for live quota monitoring without modifying Claude's internals or requiring browser extensions. The subprocess has a 150ms timeout to avoid blocking Claude's UI, and errors fail silently to prevent breaking the parent process.

Data flows through a two-phase ingestion model: live transcripts feed "provisional" records that become "archival" after 24 hours. This stabilization window solves a practical problem—Claude Code appends to transcripts as conversations continue, so immediate analysis would miss follow-up turns. By waiting a day before locking in summaries, Token Envy captures complete conversation arcs while still showing current trends in the provisional dataset. The tool also tracks exclusions explicitly: requests under 10 tokens, sessions with hour-scale gaps, and invalid timestamps all get counted separately and displayed in the UI, showing you exactly what data was filtered and why.

Gotcha

Token Envy is architecturally single-user and single-machine. SQLite storage at ~/.tokenenvy means your analytics are siloed per workstation, with no sync mechanism for multi-device workflows. If you're on a team trying to detect organization-wide performance degradation or share baseline comparisons, this doesn't help—everyone gets their own isolated metrics. There's also no documented backup/restore flow, so migrating to a new machine means losing historical data unless you manually copy the database.

The Node.js 22.13+ requirement is a hard floor that excludes LTS environments. Many enterprise teams are locked to Node 18 or 20 by central IT policies, and Token Envy won't run at all. The status-line integration is similarly fragile: it requires write access to ~/.claude/settings.json and manually running an install script. If Claude Code is running in a containerized environment, read-only filesystem, or corporate-managed configuration, this silently fails. More critically, there's no export to structured formats like CSV or Parquet, so you can't feed metrics into existing observability stacks or programmatic analysis. You're locked into the built-in visualizations and PNG share cards.

Verdict

Use if: You're a security researcher, power user, or prompt engineer running hundreds of Claude Code requests per month and need to quantify API performance shifts without shipping logs to third parties. The mix-adjusted Speed Index and effective throughput metrics are genuinely useful for detecting Anthropic infrastructure throttling versus prompt complexity changes, and the HMAC pseudonymization means you can share performance data publicly without leaking client work. This is the tool for solo practitioners doing high-volume work where privacy matters more than team collaboration. Skip if: You're a casual Claude user running fewer than 100 requests per month (the baseline requirements won't trigger), you need multi-user visibility or centralized analytics (it's architecturally local-only), or you want to integrate metrics into existing BI tools (no structured export). Also skip if you're in an enterprise environment with locked Node versions or containerized Claude deployments—the installation friction isn't worth it for light usage.