> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

FrontierHarness: The First Benchmark That Tests AI Coding Tools, Not Just AI Models

[ View on GitHub ]

FrontierHarness: The First Benchmark That Tests AI Coding Tools, Not Just AI Models

Hook

Nine different AI coding harnesses running the same model on the same tasks produced a 17.5x cost difference. Same brain, radically different monthly bills.

Context

The AI coding tool market has a measurement problem. When developers evaluate Claude Code, Codex, or OpenCode, they're actually evaluating two things at once: the underlying language model (GPT-4, Claude, etc.) and the harness—the orchestration layer that manages prompts, caching, retries, and tool calls. Existing benchmarks like SWE-bench measure end-to-end performance but can't tell you whether your 40% success rate comes from a weak model or inefficient prompt engineering. If you swap from one harness to another while keeping the same model, will your costs double? Will reliability improve? Nobody knew, because nobody was isolating the infrastructure variable.

FrontierHarness fixes this by treating coding agent harnesses as the subject of evaluation rather than the model. It's a meta-benchmark: instead of asking "which model solves the most coding tasks," it asks "which orchestration layer solves tasks most efficiently given a fixed model budget." The project locks the model (Kimi K3), locks the tasks (30 problems from Terminal-Bench and DeepSWE), and varies only the harness. The result is the first apples-to-apples comparison of Claude Code, Exo Harness, Pi Agent, and six other frameworks—showing that harness design decisions around context management and caching create massive economic impact independent of model intelligence.

Technical Insight

Per-Task Execution (360 runs)

Restore snapshot

Provision VM

Clone state

Inject stub keys

Proxy API calls

Execute markdown task

Runs in

Generates evidence

Score + normalize costs

Golden Checkpoint

VM Snapshot

Evaluation Orchestrator

Runta Runtime API

Fresh VM Instance

Egress Proxy

API Key Stub

Coding Harness

9 variants

Deterministic Verifier

Results Database

Cost + Pass/Fail

System architecture — auto-generated

The architecture centers on reproducible golden checkpoints, a technique borrowed from systems research but novel in AI evaluation. Before any tasks run, FrontierHarness provisions a Linux VM through Runta's runtime API, installs all nine harnesses (Claude Code, Codex, DeepSeek Harness, Exo Harness, OpenCode, Pi Agent, and others), then freezes that entire state as a VM snapshot. Each of the 30 tasks gets a fresh restore from that checkpoint, guaranteeing identical starting conditions. This eliminates the warm-cache bias that plagues traditional benchmarks where earlier tasks pollute the environment for later ones. With 9 harnesses × 12 model configurations × 30 tasks, that's 360 independent VM provisions—computationally expensive but methodologically bulletproof.

The evaluation workflow orchestrates everything through Runta's API. Here's the core execution loop:

// Simplified from the actual evaluation orchestrator
async function evaluateTask(harness, task, modelConfig) {
  // Restore golden checkpoint to pristine VM state
  const runtime = await runta.restoreSnapshot('golden-checkpoint');
  
  // Inject API keys through egress proxy (agent only sees stubs)
  await runtime.configureProxy({
    stubKey: 'runta-secret-stub',
    realKey: process.env.ANTHROPIC_API_KEY
  });
  
  // Execute task using harness-specific interface
  const result = await runtime.executeSkill({
    harness: harness.name,
    instruction: task.markdown,
    timeout: 3600,
    workdir: '/workspace'
  });
  
  // Collect evidence: logs, file diffs, API call traces
  const evidence = await runtime.collectEvidence([
    '/var/log/harness.log',
    '/workspace/**/*.diff',
    '/tmp/api-calls.jsonl'
  ]);
  
  // Run deterministic verifier (not LLM-as-judge)
  const passed = await task.verifier.check(evidence);
  
  // Normalize costs by repricing cache reads consistently
  const normalizedCost = repriceCacheReads(
    result.apiCalls,
    result.cacheHits
  );
  
  return { passed, cost: normalizedCost, evidence };
}

The API key injection through the egress proxy is architecturally elegant. Instead of mounting secrets as environment variables or files (which leak into logs and agent memory), the runtime only exposes a stub credential runta-secret-stub. When the harness makes API calls, the proxy intercepts outbound HTTPS, swaps the stub for real credentials, and forwards the request. The agent never sees production keys, preventing accidental credential leakage in debug output or context windows.

The skill-based execution model is particularly clever. Rather than hardcoding how to invoke each harness's CLI, FrontierHarness ships agent-neutral markdown instructions:

# Task: Fix the authentication bug

You are running inside a VM with the ${HARNESS_NAME} harness pre-installed.
The bug is in `/workspace/auth.py` line 47.
Run the test suite with `pytest tests/test_auth.py` to verify your fix.

Success criteria:
- All tests pass
- No changes outside `/workspace/auth.py`

This markdown gets fed to the harness's native interface. Claude Code interprets it through its chat UI, Codex through its prompt API, Pi Agent through its task planner. The framework effectively uses AI agents to evaluate other AI agents at scale without maintaining harness-specific integration code. When a new harness launches, maintainers can install the evaluation skill and run the benchmark without understanding FrontierHarness's internals.

The cache repricing methodology addresses a real apples-to-oranges problem. Different harnesses use wildly different prompting strategies: some frontload massive system prompts that cache well, others use minimal context and iterate. When Claude charges 10% of normal rates for cache hits, raw provider billing isn't comparable. FrontierHarness normalizes this by repricing all first-turn cache reads at a consistent rate across harnesses, using proprietary data about prompt structure. A harness that achieves 60% success with heavy caching gets cost-adjusted to reflect what that would cost with uniform cache behavior. This normalization is what reveals the true 17.5x cost spread—some harnesses are just burning tokens on redundant context.

Verifier-based scoring avoids the meta-evaluation trap. Instead of asking GPT-4 to judge whether Claude Code's output is correct (introducing another model as a variable), each task ships with deterministic checkers: does pytest pass? Does the file hash match expected output? Did the API return 200? This makes the benchmark reproducible across runs and auditable by humans. The 30 tasks include terminal automation challenges ("write a script that monitors CPU and sends Slack alerts") and software engineering problems ("fix the race condition in the webhook handler"), verified by unit tests and integration checks.

Gotcha

The biggest limitation is single-model lock-in. FrontierHarness exclusively uses Kimi K3, so you can't determine whether certain harnesses pair better with specific models. Maybe Claude Code's prompt engineering extracts more from Claude Opus than from Kimi, while Exo Harness optimizes for GPT-4's quirks. The benchmark can't answer that question—it only tells you how harnesses perform with one brain. For teams standardizing on a specific model this is fine, but for organizations evaluating model flexibility, it's a blind spot.

Task coverage is narrow: 30 problems from just two suites (Terminal-Bench and DeepSWE) doesn't represent real-world diversity. There are no web scraping scenarios, no data science notebooks, no DevOps automation, no multi-repository refactorings. The benchmark skews toward terminal automation and isolated bug fixes. If your production workload is "build a web scraper that handles CAPTCHA," these results won't predict harness performance.

The cache repricing methodology is opaque—FrontierHarness mentions "data that is not public" used for normalization. Third parties can't verify the cost calculations or apply the same normalization to new harnesses. You have to trust their repricing model, which is reasonable for a first-generation benchmark but limits scientific reproducibility. Open-sourcing the repricing algorithm would let the community validate and improve it.

Infrastructure requirements are steep. Reproducing the benchmark requires a Runta runtime subscription, 4-8 vCPUs per task, 50GB disk per VM, and custom orchestration logic. You can't run this on GitHub Actions or typical CI/CD pipelines. The published results are immediately useful, but creating new benchmarks or adding tasks requires non-trivial investment in Runta infrastructure. This isn't a "clone and npm test" repository.

Verdict

Use if: you're selecting a coding agent harness for production workloads where cost predictability matters, you're building a new harness and need competitive intelligence on cache utilization and prompt efficiency, or you're a decision-maker who needs quantitative evidence that infrastructure choices (not just model selection) drive ROI. The published results are immediately actionable—if you're paying for Claude Code but Exo Harness achieves similar pass rates at 5x lower cost, that's a migration case. Skip if: you need to compare models rather than harnesses (use SWE-bench directly for that), your tasks involve domains not covered by Terminal-Bench and DeepSWE (web scraping, data science, DevOps), you lack the infrastructure to run isolated VM-based evaluations, or you need to understand how harnesses perform across multiple models. The real insight here is proving that agent orchestration is an independent optimization surface with measurable business impact—harness choice isn't just a developer preference issue, it's a cost center.