AutoAgent: Teaching AI Agents to Rewrite Themselves for Better Benchmark Scores
Hook
Most AI agent frameworks ask you to write prompts and tools. AutoAgent asks: what if the agent rewrote its own source code overnight to score higher on your benchmarks?
Context
Building effective AI agents is expensive. You write a system prompt, define tools, test against benchmarks like GAIA or SWE-bench, discover your agent fails at function calling, tweak the prompt, run tests again, repeat. Each iteration burns hours of developer time and API credits. Frameworks like LangChain or AutoGen help with scaffolding, but the core loop—diagnose failure, modify code, retest—remains manual.
AutoAgent flips this paradigm. Instead of you modifying agent code, a meta-agent does it. You provide high-level instructions in program.md, point it at Harbor-formatted benchmarks, and let it run overnight. The meta-agent reads the current harness (agent.py), executes it against test tasks in isolated Docker containers, checks scores in /logs/reward.txt, then edits agent.py to improve performance. Keep changes that increase total score, discard the rest. This is hill-climbing applied to code: no gradients, no reinforcement learning, just greedy optimization on a single metric. It trades interpretability and control for autonomous iteration, letting you wake up to an agent tuned for your specific benchmark suite.
Technical Insight
The architecture rests on three pillars: a single-file harness, nested agent execution, and a fixed adapter boundary. Understanding how these interact reveals both the system's power and constraints.
Single-File Registry Pattern
Typical agent frameworks spread logic across modules—prompts in one file, tools in another, routing elsewhere. AutoAgent collapses everything into agent.py using registration decorators:
@agent.register_system_prompt
def system_prompt():
return """
You are an autonomous agent solving complex tasks.
Use tools methodically. Break problems into steps.
Always verify your work before concluding.
"""
@agent.register_tool
def search_web(query: str) -> str:
"""Search the internet for information."""
# Implementation here
return results
@agent.register_tool
def execute_python(code: str) -> str:
"""Run Python code in a sandbox."""
# Sandboxed execution
return output
# Harbor adapter boundary - DO NOT MODIFY
def run_task(task_description: str) -> dict:
"""Fixed interface for benchmark execution."""
result = agent.execute(task_description)
return {"answer": result, "metadata": agent.get_trace()}
This registry pattern gives the meta-agent a predictable edit surface. Want to improve tool descriptions? Modify the docstring. Need better routing? Change the system prompt. Everything lives in one file, so the meta-agent doesn't need to track imports or resolve dependencies across modules. The tradeoff is obvious: agent.py becomes unwieldy beyond ~500 lines, and modular patterns like plugin systems break the optimization loop.
Nested Agent Execution Model
The meta-agent and inner agent operate at different abstraction levels. The meta-agent uses a coding-focused LLM (typically Claude 3.5 Sonnet or GPT-4) to read program.md instructions like "improve function calling accuracy" or "add better error handling for API timeouts." It analyzes the current agent.py, reviews logs from failed tasks, and generates a diff:
# Meta-agent reasoning (conceptual)
def optimize_iteration():
current_code = read_file("agent.py")
instructions = read_file("program.md")
previous_logs = read_file("results.tsv")
diagnosis = meta_llm.analyze(
code=current_code,
instructions=instructions,
failures=previous_logs
)
proposed_diff = meta_llm.generate_improvements(
diagnosis=diagnosis,
constraints=["preserve Harbor adapter", "keep under 200k tokens"]
)
apply_diff(proposed_diff, "agent.py")
new_score = run_benchmarks()
if new_score > previous_score:
commit_changes()
else:
revert()
The inner agent just executes tasks. It doesn't know it's being optimized. When agent.py runs against a benchmark task, it reads the task description, calls tools, and produces an answer. Docker isolation ensures it can't escape the sandbox or tamper with the meta-agent's logic.
Harbor Adapter Boundary
The run_task() function is marked immutable. This boundary prevents the meta-agent from gaming benchmarks—no modifying score calculations, no sneaking peeks at ground truth answers, no changing how tasks are loaded. The meta-agent can only improve the agent's problem-solving logic, not the evaluation mechanism. This trust boundary is critical: without it, the optimization would degenerate into exploiting eval bugs rather than solving tasks.
Docker Isolation Per Task
Each benchmark task runs in its own container inheriting from autoagent-base:
# Base image with agent harness
FROM autoagent-base:latest
# Task-specific environment
RUN pip install selenium beautifulsoup4
COPY task_files/ /workspace/
# Entry point executes test.sh
CMD ["/workspace/test.sh"]
The test.sh script invokes agent.py with the task description, captures output, and writes a score to /logs/reward.txt. Deterministic tasks use exact string matching; open-ended tasks use an LLM judge. This containerized approach enables parallel execution—run 50 tasks simultaneously without interference—and makes benchmarks portable. You can share Harbor task definitions across different agent implementations.
The Hill-Climbing Loop
AutoAgent doesn't use gradients or policy optimization. It's pure greedy search: try a change, measure aggregate score across all tasks, keep if better. This simplicity is both strength and weakness. No complex training infrastructure, no hyperparameter tuning, just edit-test-accept/reject. But it gets stuck in local maxima. An improvement on five tasks that regresses catastrophically on one might still increase total score, hiding the regression. And there's no backtracking—if the agent goes down a dead end, it can only move forward from there.
Gotcha
The single-file constraint creates a brutal scaling ceiling. Once agent.py exceeds a few hundred lines, it becomes cognitively overwhelming for the meta-agent to reason about. Complex changes like "add a planning phase with tree search" or "implement RAG with vector stores" start producing buggy diffs because the meta-agent loses track of code structure. There's no clear path to multi-file refactoring without breaking the optimization loop.
Score aggregation hides critical failures. If your benchmark has 100 tasks and the agent improves from 60% to 65% overall, you'd celebrate. But dig into per-task results and you might find it now fails completely on basic tasks while solving a few hard ones—trading robustness for local optimization. AutoAgent's greedy hill-climbing doesn't do Pareto optimization or maintain performance profiles. You get a single number that trends upward, with no insight into whether improvements are generalizable or brittle overfitting to your specific benchmark suite.
The meta-agent has no safety rails beyond the adapter boundary. It can introduce syntax errors, infinite loops, or race conditions that only trigger under specific task conditions. You might wake up to an agent that scores 5% higher but occasionally deadlocks, or one that works on the benchmarks but crashes in production because it makes assumptions about task structure. There's no static analysis, type checking, or semantic validation—just LLM-generated code changes accepted on score delta alone.
Verdict
Use if: you're running closed-loop optimization experiments on well-defined benchmarks (GAIA, WebArena, SWE-bench), need rapid iteration without manual tweaking every prompt, and accept that interpretability matters less than raw scores. It's brilliant for overnight optimization runs that would otherwise burn weeks of engineering time, especially in research contexts where benchmark performance is the primary goal. Use if: you have a stable benchmark suite, can tolerate occasional broken iterations, and plan to extract successful patterns into a production framework afterward. Skip if: you need to understand why changes improve performance, require multi-objective optimization (accuracy + latency + cost), plan to deploy agents in production where robustness trumps benchmark scores, or want human-in-the-loop development with fine-grained control. Skip if: your agent logic exceeds 500 lines or requires modular architecture with plugins and extensions. AutoAgent is a research tool for bootstrapping agent designs, not a production framework for deploying reliable systems.