> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

Switchyard: Composable LLM Routing Without the Gateway Lock-In

[ View on GitHub ]

Switchyard: Composable LLM Routing Without the Gateway Lock-In

Hook

Most LLM proxies force you to replace your entire API gateway just to route between GPT-4 and Claude. Switchyard inverts this: it's a library that yields control back to your code for every model call, making routing logic composable rather than all-or-nothing.

Context

If you're running LLM applications in production, you've hit the model selection problem: GPT-4 is expensive and slow for simple queries, but Llama hallucinates on complex reasoning. The obvious solution is routing—send easy requests to cheap models, hard ones to frontier models. But existing proxies like LiteLLM and Portkey are monolithic: they want to own your HTTP layer, authentication, rate limiting, and observability stack.

This creates a dilemma for teams with existing infrastructure. You already have an API gateway handling auth, a service mesh for traffic management, and Prometheus for metrics. Adopting a smart LLM proxy means either running it as a sidecar (doubling network hops) or ripping out your gateway and migrating everything to the proxy's ecosystem. Switchyard solves this by separating routing decisions from HTTP execution. It's a Rust library that tells you which model to call, then hands control back so you can make the request using your existing client code, connection pools, and retry logic.

Technical Insight

Caller-Owned

Core Library

OpenAI/Anthropic

format

Parsed request

Target selection

no HTTP call

Request

Backend-native

format

Response

Normalized

response

Process response

Continue?

Escalate?

Final response

Client Request

switchyard-server

Protocol Detection

switchyard-libsy

Routing Algorithms

Protocol Translation

Layer

Caller's HTTP

Client Code

LLM Backend

APIs

System architecture — auto-generated

The architectural innovation is callback-driven control flow. Unlike traditional proxies that accept a request, route it, and return a response, Switchyard's core library (switchyard-libsy) never makes HTTP calls itself. Instead, routing algorithms return target selections, and your code invokes the models:

use switchyard_libsy::algorithm::stage::{StageRouter, StageConfig};
use switchyard_libsy::types::{Request, ModelResponse};

// Configure a stage router: try fast tier first, escalate on tool errors
let config = StageConfig {
    stages: vec![
        Stage { models: vec!["llama-3.1-8b"], max_retries: 1 },
        Stage { models: vec!["gpt-4o"], max_retries: 0 },
    ],
    escalation_signals: vec!["tool_error", "json_parse_failure"],
};
let mut router = StageRouter::new(config);

// Router yields model targets, doesn't call them
let decision = router.route(&request)?;
for target in decision.targets {
    let response = your_llm_client.call(&target.model_id, &request).await?;
    
    // Router inspects response to decide if we escalate
    let should_continue = router.process_response(&response)?;
    if !should_continue { break; }
}

This callback pattern means Switchyard integrates into existing codebases without replacing the transport layer. If you already have an Axum server with OAuth middleware and connection pooling to OpenAI, you keep all of that—Switchyard just adds routing logic between your auth layer and your HTTP client.

The protocol translation layer is where Rust's type system shines. Switchyard defines traits for OpenAIChatRequest, AnthropicMessagesRequest, and unified response types. The translation is bidirectional and streaming-aware:

// Client sends OpenAI format, backend is Anthropic
let openai_req: OpenAIChatRequest = parse_from_client(body)?;
let anthropic_req = translate_to_anthropic(&openai_req)?;

// Make the call (using YOUR http client)
let anthropic_stream = http_client.post(anthropic_url)
    .json(&anthropic_req)
    .send()
    .await?;

// Translate SSE events back to OpenAI format without buffering
let openai_stream = anthropic_stream
    .map(|event| translate_event_to_openai(event));

The streaming translation is zero-copy for message content—only protocol metadata gets rewritten. This matters for long-form generation where buffering the entire response would spike memory and add seconds to time-to-first-token.

Routing algorithms are stateful and conversation-aware. The stage router tracks tool invocation results across turns, so if your agent tries to call read_file() and gets a permission error, that's a routing signal to escalate to a smarter model for the next turn—no additional LLM classifier call needed:

impl StageRouter {
    fn should_escalate(&self, history: &ConversationHistory) -> bool {
        history.last_turn()
            .tool_results
            .iter()
            .any(|r| r.is_error && self.config.escalation_signals.contains("tool_error"))
    }
}

This is cheaper than running a classifier LLM on every request (the approach LiteLLM's dynamic routing takes), because most escalation patterns are deterministic once you track conversation state. For complex cases, Switchyard does support LLM-based classification, but it's opt-in rather than required.

The escalation router implements speculative execution at the LLM level: it runs a weak model unconditionally, then uses a judge LLM to evaluate response quality. If the judge rejects it, Switchyard re-routes to a strong model and returns that response instead. The judge can be a small, fast model (even a fine-tuned distilBERT for specific domains), so the overhead is controllable. The Prometheus metrics expose routing overhead separately from model latency, letting you measure whether smart routing actually improves cost-per-quality in your workload.

Gotcha

The pre-alpha maturity is real. API surface is unstable, documentation is sparse outside of code comments, and the TOML configuration format has breaking changes between commits. If you need this today, you're forking and maintaining a pin—treat it as reading group material for understanding routing architectures, not a drop-in dependency.

Protocol translation is fundamentally lossy. Anthropic's thinking blocks (the <thinking> XML tags Claude uses for chain-of-thought) don't map to OpenAI's function calling format. OpenAI's tool_choice: required parameter forces tool use, but Anthropic has no equivalent. If you're building agentic systems that rely on provider-specific features—like extracting Claude's reasoning steps or using GPT-4's strict JSON schema mode—round-tripping through Switchyard will break those workflows. The escalation router doubles token costs and latency for escalated requests because the weak tier always runs first. There's no early classification to skip the weak model if the request is obviously hard ("write a Rust compiler" shouldn't hit Llama before GPT-4). You could implement this yourself using the callback architecture, but it's not built-in.

Verdict

Use Switchyard if you're building LLM infrastructure as a platform team and need routing logic you can audit, extend, and version control as code—especially if you already have API gateway components and want to add intelligence without vendor lock-in. It's ideal for research infrastructure (benchmarking prompt injection defenses across models), agentic platforms (routing tool failures to stronger models), or security teams running red-team evaluations where you need to control every layer of the stack. The callback architecture and Rust typing make this a library you integrate, not a service you deploy. Skip if you need production-ready multi-tenancy, authentication, or billing—LiteLLM and Portkey handle those out of the box. Also skip if you're a small team without Rust expertise or your routing needs are just weighted round-robin (use Envoy). The sweet spot is offensive security orgs, AI research labs, or platform engineering teams at companies building LLM products who want routing as a composable primitive rather than a walled garden.