> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

JIT-Agent: When Your AI Writes Its Own Agent Architecture

[ View on GitHub ]

JIT-Agent: When Your AI Writes Its Own Agent Architecture

Hook

What if instead of hand-coding agent architectures, you let an LLM generate task-specific scaffolds on-the-fly? JIT-Agent treats agent design itself as a learnable capability that improves without retraining the base model.

Context

Every agent framework forces you to make upfront architectural decisions: Will you use ReAct loops or planning-first approaches? Short-term memory or full conversation history? Tool orchestration or capability graphs? These choices are baked into LangGraph nodes, AutoGen conversation patterns, and CrewAI role definitions. But here's the problem: the optimal scaffold varies dramatically by task class. A research QA agent benefits from hierarchical planning and long-term memory. A code execution agent needs tight action loops with minimal planning overhead. A travel booking agent requires stateful multi-step orchestration.

Most frameworks solve this with configuration — you write YAML or builder patterns to customize the architecture. JIT-Agent takes a fundamentally different approach: it uses an LLM to generate the entire agent scaffold as Python code at runtime, specialized for your specific task. The meta-model receives your task specification, a tool registry, and examples of prior harnesses, then emits a complete agent implementation that adheres to standardized interfaces. The system maintains an evolving archive of successful harnesses, improving through test-time revision while keeping the generator model frozen. It's meta-programming for agent systems: the code that runs your agent is itself written by an agent.

Technical Insight

JIT-Agent's architecture splits agent development into three distinct model roles. The meta-model (typically a 27B parameter checkpoint or GPT-4) writes harness code. The execution model (can be as small as 7B) runs the generated agent loop. The judge model evaluates output quality. This separation means you can use an expensive, capable model to generate sophisticated architectures that cheaper models execute — amortizing design cost across many runs.

The core abstraction is HarnessFactory, which defines four standardized modules every harness must implement:

class HarnessFactory:
    def create_memory(self) -> Memory:
        # Returns memory system (conversation history, facts, etc)
        pass
    
    def create_planner(self, memory: Memory) -> Planner:
        # Returns planning strategy (hierarchical, reactive, none)
        pass
    
    def create_action_executor(self) -> ActionExecutor:
        # Returns tool/API execution logic
        pass
    
    def create_orchestrator(self, memory, planner, executor) -> Orchestrator:
        # Returns main agent loop coordinating the above
        pass

When you submit a task, the meta-model generates multiple candidate implementations of these interfaces. Here's a simplified example of what the generator might produce for a research-heavy task:

class ResearchHarness(HarnessFactory):
    def create_memory(self):
        # Hierarchical memory for complex research
        return HierarchicalMemory(
            short_term_capacity=10,
            long_term_indexer="semantic"
        )
    
    def create_planner(self, memory):
        # Decompose research into sub-questions
        return HierarchicalPlanner(
            decomposition_depth=3,
            memory=memory
        )
    
    def create_action_executor(self):
        return ToolExecutor(
            retry_logic="exponential_backoff",
            parallel_execution=True
        )
    
    def create_orchestrator(self, memory, planner, executor):
        return PlanExecuteOrchestrator(
            planning_threshold=5,  # Replan every 5 actions
            memory=memory,
            planner=planner,
            executor=executor
        )

For a simple code execution task, the same meta-model might generate a radically different harness:

class CodeExecHarness(HarnessFactory):
    def create_memory(self):
        # Minimal memory for stateless execution
        return BufferMemory(capacity=3)
    
    def create_planner(self, memory):
        # No planning, just react
        return NullPlanner()
    
    def create_action_executor(self):
        return ToolExecutor(
            retry_logic="fail_fast",
            parallel_execution=False
        )
    
    def create_orchestrator(self, memory, planner, executor):
        return ReActOrchestrator(
            max_iterations=10,
            memory=memory,
            executor=executor
        )

The generation process creates N candidates (typically 3-5), then a selector chooses the best. JIT-Agent supports two selection strategies. The judge-based selector uses an LLM to evaluate each candidate against the task requirements — accurate but expensive. The log-probability selector is more interesting: it uses the meta-model's own token-level confidence scores as a quality signal. Harnesses the model generates with high certainty tend to be more reliable:

# Simplified selector logic
def select_harness_by_logprob(candidates, meta_model):
    scores = []
    for candidate in candidates:
        # Tokenize the generated code
        tokens = meta_model.tokenize(candidate.code)
        # Get log probabilities from generation
        logprobs = meta_model.get_logprobs(tokens)
        # Average log probability as quality signal
        scores.append(sum(logprobs) / len(logprobs))
    
    return candidates[argmax(scores)]

This approach requires local tokenizer access but avoids expensive judge calls. The system tracks generated harnesses in an archive, enabling test-time evolution. When a harness succeeds, it becomes training data for future generations. When it fails, the meta-model can revise it based on execution traces. Crucially, the meta-model itself stays frozen — improvements come from better scaffolds in the archive, not weight updates.

The benchmark adapters in benchmark/ demonstrate the power of protocol-agnostic design. Each adapter defines task I/O shape and scoring, but the harness handles agent structure. A DeepSearchQA adapter provides research questions and measures answer accuracy. A TravelPlanner adapter gives booking constraints and checks itinerary validity. The same harness selection and execution pipeline works for both, because the harness code adapts to the task protocol rather than the framework enforcing a single architecture.

Gotcha

JIT-Agent's three-model architecture creates operational complexity that matters in production. A single benchmark run with three candidate harnesses requires three meta-model generation calls, potentially three judge model evaluation calls, and then the execution model runs the selected harness. There's no batching across pipeline stages, and if you're using API-based models, costs scale quickly. The project provides no built-in rate limiting or quota management — if your Serper API key hits limits halfway through a benchmark suite, the run fails with no automatic retry logic.

The log-probability selector sounds elegant but introduces subtle brittleness. It requires downloading the full meta-model tokenizer and assumes the tokenizer API remains stable. If you quantize the model, distill it, or the checkpoint format changes, the selector breaks. The codebase provides no fallback mechanism. Judge-based selection works with any API endpoint but the documentation explicitly warns it's slower and more expensive — you're choosing between operational fragility and cost.

The repository includes a 27B parameter checkpoint for evaluation, but zero training code. The paper describes training data construction from harness execution traces and self-revision loops, but none of that machinery is released. You can run inference with the provided model, but you cannot reproduce the training process or adapt the approach to your own task distributions. For a research project claiming harness learning as a core contribution, the absence of training infrastructure is a significant gap. You're effectively getting a demo of what's possible, not a tool for doing it yourself.

Verdict

Use JIT-Agent if: You're researching agent architecture search and want to explore meta-programming approaches to scaffold design. You're operating complex multi-model pipelines already and have infrastructure for vLLM/SGLang deployment. You have narrow, high-value task domains (specialized research, multi-step planning) where harness specialization could justify the operational overhead. You're comfortable reading research code and extending it.

Skip if: You need a production-ready agent framework with predictable costs and operational simplicity — LangGraph or AutoGen will be far easier to deploy and debug. You're working in languages other than Python — the approach is tightly coupled to Python code generation and execution. You want to reproduce the training process or adapt the meta-model to your own tasks — only inference code is provided. You need general-purpose agents that handle diverse tasks — JIT-Agent's value proposition is task-specific specialization, which requires multiple runs to amortize generation costs. The project demonstrates that agent scaffolding can be learned and specialized, which is intellectually compelling. But it's research infrastructure, not a framework you'd choose for building production agents today.