> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

AX: Google's Agentic Orchestrator Treats AI Agents Like Kubernetes Workloads

[ View on GitHub ]

AX: Google's Agentic Orchestrator Treats AI Agents Like Kubernetes Workloads

Hook

Google built AX because standard Kubernetes pods weren't designed for workloads that burn hundreds of dollars in inference loops, need to pause mid-execution to save costs, and require network isolation before executing LLM-generated code against your production APIs.

Context

If you've tried running autonomous agents at scale, you've hit the orchestration wall. Agents aren't stateless services that respond to requests and die. They're not batch jobs that run once and finish. They're stateful, long-running processes that loop indefinitely, make unpredictable API calls, and consume expensive LLM inference tokens while waiting for external events. Standard Kubernetes treats them like any other container: deploy a pod, let it run, restart on failure. But this model breaks when an agent enters an infinite reasoning loop and racks up $500 in OpenAI charges overnight, or when you need to pause 10,000 agents to swap models without losing their conversation state.

AX is Google's answer to this orchestration gap. Built on top of Agent Substrate (Google's sandboxed actor runtime), AX provides a declarative control plane specifically for agentic workloads. It introduces purpose-built primitives: Tasks that describe agent execution, Workspaces that pre-wire Git repos and tool servers before boot, Gateways that allowlist network egress at the sandbox level, and Models that inject LLM credentials without baking secrets into images. The architecture recognizes that agents are fundamentally different workloads requiring explicit lifecycle controls, dependency composition, and security boundaries that standard container orchestrators never anticipated.

Technical Insight

AX's core abstraction is the Task manifest, which looks deliberately Kubernetes-like but introduces agent-specific fields. Here's what deploying an agent actually looks like:

apiVersion: ax.dev/v1alpha1
kind: Task
metadata:
  name: customer-support-agent
spec:
  workspace:
    repos:
      - name: support-kb
        url: https://github.com/company/support-knowledge
        ref: main
        mountPath: /workspace/kb
    mcpServers:
      - name: zendesk-tools
        image: company/zendesk-mcp:v2
        endpoint: unix:///var/run/mcp/zendesk.sock
  model:
    provider: openai
    name: gpt-4
    credentialsSecret: openai-prod-key
  gateway:
    egress:
      allowList:
        - zendesk.com
        - api.openai.com
        - company-internal.corp
  resources:
    cpuMillis: 1000
    memoryMB: 2048
  runner:
    image: company/python-agent-runner:v1
    command: ["python", "/app/agent.py"]

When you ax apply this manifest, the controller performs a reconciliation sequence that's more sophisticated than standard pod scheduling. First, it resolves the Workspace by cloning Git repos into a shared volume and provisioning MCP (Model Context Protocol) servers as sidecar processes with Unix socket endpoints. This means your agent boots into an environment where tools are already wired and knowledge bases already mounted—no cold-start setup phase burning tokens.

Second, the Gateway egress rules get translated into network namespace filters within the Agent Substrate sandbox. This isn't Kubernetes NetworkPolicy (which operates at pod-to-pod level); it's process-level egress filtering enforced by the sandbox runtime. When your agent's LLM generates code that tries to curl an unauthorized domain, the request dies at the network namespace boundary before touching the cluster network. This is critical security architecture: you're running code generated by a probabilistic model that might hallucinate API calls to AWS or attempt to exfiltrate data.

The clever part is how Model credentials get injected. AX runs a metadata server inside each sandbox (borrowed from the EC2 IMDS pattern) that the runner can query:

# Inside your agent runtime
import requests
import os

metadata_url = os.getenv('AX_METADATA_SERVER')
model_config = requests.get(f'{metadata_url}/model').json()
# Returns: {"provider": "openai", "model": "gpt-4", "apiKey": "sk-..."}

client = OpenAI(api_key=model_config['apiKey'])

This decouples credentials from container images and ConfigMaps. You can rotate API keys, swap model providers, or change models without rebuilding images—the metadata server injects current config at runtime. For teams running thousands of agents across development, staging, and production, this separation of concerns is operationally critical.

The suspend/resume primitive addresses the cost control problem directly. Unlike Kubernetes pods that either run or terminate, AX Tasks support explicit lifecycle transitions:

# Agent is burning tokens in a loop waiting for user input
ax suspend customer-support-agent
# Checkpoints actor state, stops execution, releases resources

# User responds 6 hours later
ax resume customer-support-agent
# Restores checkpoint, continues from exact state

Under the hood, this relies on Agent Substrate's actor model checkpointing the agent's memory, conversation history, and internal state to Redis. When resumed, the runner container restarts and the metadata server provides a restore token that hydrates the previous state. This is fundamentally different from Kubernetes StatefulSets, which preserve disk volumes but reboot processes from scratch. AX preserves in-memory execution state across suspend cycles, letting you pause expensive inference workloads during idle periods without losing context.

The standardized runner contract is what makes this architecture extensible. Your runner image just needs to implement the AX init protocol: query the metadata server for workspace topology and model config, hydrate checkpoints if resuming, execute agent logic, and expose health/metrics endpoints on expected ports. The control plane doesn't care whether you're running LangChain, AutoGPT, custom Python, or compiled Go agents. This decoupling mirrors how Kubernetes doesn't care what's inside your container as long as it responds to signals and exposes ports correctly. You can evolve agent frameworks without touching orchestration infrastructure.

Gotcha

AX's biggest limitation is its alpha maturity and hard dependency on Agent Substrate. The v1alpha1 API is explicitly unstable—the maintainers warn that schema changes will break existing manifests. For production use, this means accepting migration work with every release and maintaining version-pinned manifests with upgrade runbooks. Worse, Agent Substrate itself is nascent technology. If you adopt AX, you're betting on two unproven projects maintaining compatibility as both evolve. A breaking change in Agent Substrate's checkpoint format could cascade into AX state corruption with no clear rollback path documented.

The Gateway egress allowlist is surprisingly rigid for infrastructure claiming to support billions of tasks. It only accepts hostnames, with no visible support for IP CIDR blocks or dynamic service discovery. If your agents need to call AWS services (which resolve to thousands of rotating IPs behind load balancers), you're stuck allowlisting *.amazonaws.com and losing fine-grained controls. There's also no mention of multi-tenancy primitives beyond basic Namespace separation—no quota enforcement, priority classes, or isolation guarantees to prevent one team's runaway agents from starving cluster resources. The documentation shows single-task examples but provides zero guidance on autoscaling policies, backpressure handling, or task prioritization when the cluster saturates. For a system claiming 'billions of tasks per cluster' scale, these are conspicuous documentation gaps that suggest the operational story isn't fully baked yet.

Verdict

Use if: You're running massive agent fleets (hundreds to thousands of concurrent agents) where lifecycle control, cost optimization, and security isolation justify bleeding-edge tooling. AX is purpose-built for teams at Google scale who need declarative workspace composition, network sandboxing, and suspend/resume primitives that standard Kubernetes genuinely lacks. It's ideal if you're already committed to Agent Substrate or building an internal agentic platform where betting on unstable APIs is acceptable. The Kubernetes-native UX (kubectx integration, YAML manifests, familiar CLI patterns) makes it compelling for platform teams who want agent orchestration that feels native to existing k8s workflows. Skip if: You're prototyping with small agent counts (under 100), need API stability in the next 6-12 months, or require proven multi-tenancy for production SaaS deployments. The alpha warnings are genuine—expect breaking changes and missing operational features. Also skip if your agents must call services behind dynamic cloud infrastructure; the hostname-only egress allowlisting is too rigid. Consider Temporal.io for durable workflow orchestration with mature failure handling, or Modal/Beam for serverless ML workloads with better DX, if you don't specifically need AX's sandboxing and workspace primitives. Remember you're adopting an entire stack (AX + Agent Substrate), not just swapping schedulers.