> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

Harness Engineering: The Discipline AI Teams Are Building Without Knowing It

[ View on GitHub ]

Harness Engineering: The Discipline AI Teams Are Building Without Knowing It

Hook

Every production AI agent you've deployed is mostly scaffolding code that has nothing to do with the model. OpenAI estimates 80% of agent reliability comes from harness design, not prompt quality—yet no one calls it 'harness engineering' until now.

Context

If you've built an AI agent in the last 18 months, you've written the same scaffolding everyone else writes: retry loops when the LLM returns malformed JSON, tool selection logic when it picks the wrong function, memory systems because context windows overflow, and permission checks because you can't let GPT-4 delete production databases. This code isn't prompt engineering and it isn't machine learning—it's infrastructure that compensates for what models can't do reliably.

The ai-boost/awesome-harness-engineering repository names this work for the first time. Created as a curated awesome-list with 3,600+ stars, it defines 'harness engineering' as the discipline of building scaffolding around AI agents: the loops, guardrails, memory systems, and orchestration layers that turn a language model into something you can deploy. The repository curates ~50 foundational essays from OpenAI, Anthropic, Google, Meta, and Microsoft into a nine-primitive taxonomy (agent loops, planning, context delivery, tool design, permissions, memory, orchestration, verification, observability). It's not code—it's a knowledge graph that legitimizes harness work as distinct from prompt engineering or ML ops, aimed at engineering leaders deciding whether to staff a harness team.

Technical Insight

The repository's core contribution is taxonomizing harness components as design primitives, each encoding an assumption about model limitations. The 'agent loop' primitive exists because models can't self-correct without structured retry logic. The 'permissions' primitive exists because models can't reason about authorization boundaries. This framing creates a strategic roadmap: as models improve, harness components become obsolete.

Consider the filesystem-based context engineering pattern documented in Microsoft's Azure SRE Agent case study (one of the curated essays). Traditional agent harnesses proliferate bespoke tools—a 'query_database' function, a 'read_config' function, a 'search_logs' function. Microsoft's SRE team built 100+ such tools and hit 45% 'Intent Met' rates. They rewrote the harness to expose everything as files:

# Traditional tool-heavy harness
tools = [
    read_database_schema(),
    query_recent_incidents(),
    get_runbook(incident_id),
    search_documentation(query)
]

# Filesystem-based harness
# Agent sees a mounted directory:
# /context/
#   schema.sql
#   recent_incidents.json
#   runbooks/
#   docs/
# All accessible via standard file operations

This architectural shift—treating context as filesystem layout rather than API surface—raised 'Intent Met' from 45% to 75%. The harness complexity collapsed because the model could use native file-reading capabilities instead of learning tool signatures. The pattern generalizes: harness engineers should minimize tool proliferation and maximize reuse of model pretraining (models were trained on filesystems, not your custom tool DSL).

The repository also documents the 2024-2026 shift from single-agent frameworks to multi-agent orchestration. The curated essays from Anthropic and OpenAI reveal that production systems now route tasks to specialist agents rather than building one god-agent. A code generation harness might look like:

class CodegenHarness:
    def __init__(self):
        self.planner = Agent(model="gpt-4", role="architect")
        self.implementer = Agent(model="claude-3.5", role="coder")
        self.reviewer = Agent(model="gpt-4o", role="security")
    
    def generate(self, spec):
        # Planner creates task breakdown
        plan = self.planner.run(spec, context=self.filesystem)
        
        # Implementer writes code per-task
        code = [self.implementer.run(task) for task in plan.tasks]
        
        # Reviewer checks for vulnerabilities
        review = self.reviewer.run(code, rules=self.security_policy)
        
        # Harness orchestrates, agents don't know about each other
        return self.merge(code) if review.approved else self.retry()

This is harness-level orchestration, not agent-level. The agents don't coordinate—the harness routes outputs. The 'planning' primitive lives in one agent, 'implementation' in another, 'verification' in a third. Benchmarks cited in the repository show 5+ percentage point swings from orchestration changes alone, independent of model quality.

The most provocative pattern comes from Martin Fowler's 'Harness Engineering' essay (heavily cited in the repository): 'humans on the loop' versus 'humans in the loop'. Traditional AI safety puts humans in the agent's execution path to approve actions. Harness engineering puts humans on the loop as engineers who design the environment—the tools, permissions, and observability—then let agents run unsupervised. The harness becomes the safety layer:

class ToolHarness:
    def __init__(self, agent):
        self.agent = agent
        self.allowed_tools = self.load_permissions(agent.role)
    
    def execute(self, tool_call):
        # Harness enforces, agent doesn't see denial
        if tool_call.name not in self.allowed_tools:
            return ToolResult(error="Permission denied", 
                            logged_to_security_audit=True)
        
        # Harness wraps with observability
        with self.trace(tool_call):
            return self.sandbox.run(tool_call)

The agent never knows permission checks exist. The harness maintains the invariant that no unauthorized tool can execute. This inverts the traditional prompt-based safety approach ('You are not allowed to delete files') into architectural enforcement. The repository's taxonomy makes this pattern nameable and citable—previously it was just 'something we hacked together'.

Gotcha

The repository's fatal limitation is that it contains zero code. Despite a 'Templates' section header in the README, you get links to Medium posts, not starter harnesses. If you're an engineer trying to ship an agent this quarter, this list wastes your time—you need LangChain's agent tutorials or Anthropic's building-agents guide, both of which provide working examples.

The curation bias is also severe. Thirteen of twenty-eight foundational links are frontier lab blog posts (OpenAI, Anthropic, Google, Meta, Microsoft). Missing entirely: open-source harness architectures from AutoGPT, the CrewAI orchestration patterns that shaped early multi-agent adoption, and LangChain's agent executor design (love it or hate it, it's the de facto harness architecture for 100k+ deployments). The list reads like it was compiled by someone who only follows corporate AI labs, not the open-source ecosystem where most harness engineering actually happens. Additionally, the nine-primitive taxonomy is asserted without justification—no explanation for why these nine categories are necessary and sufficient, or how they map to existing software patterns like middleware or event loops.

Verdict

Use this repository if you're an engineering leader justifying headcount for an 'agent infrastructure' or 'harness engineering' team. The curated essays provide authoritative citations from OpenAI and Anthropic to legitimize scaffolding work as a distinct discipline, not just 'glue code.' It's valuable for strategic planning—deciding whether to invest in harness tooling versus buying better models. The taxonomy also helps in roadmap discussions: if you name 'memory' and 'orchestration' as separate primitives, you can staff them separately. Skip this if you're an IC engineer building an agent—you'll spend two hours reading essays when you need code samples and API references. Go straight to Anthropic's 'Building Effective Agents' guide (one comprehensive resource) or LangChain's docs (working examples). Also skip if you need security perspectives—the list has no coverage of adversarial harness attacks or tool-use injection despite a 'Security & Permissions' section. This is a reading list for architects, not a toolbox for builders.