> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

OpenUI: A Custom DSL That Cuts LLM Token Costs by 60% (At the Price of Everything Else)

[ View on GitHub ]

OpenUI: A Custom DSL That Cuts LLM Token Costs by 60% (At the Price of Everything Else)

Hook

What if you could cut your OpenAI bill in half by teaching GPT-4 a new language—one that only your application understands? OpenUI does exactly this, and the engineering trade-offs are fascinating.

Context

Every time an LLM generates UI, you're paying for tokens. When Vercel AI SDK streams a button component as JSON, you're burning tokens on quotes, braces, repeated keys, and commas. A simple form might consume 2,000 tokens as {"type": "button", "props": {"variant": "primary", "children": "Submit"}} when the actual semantic content is maybe 200 tokens. At scale—think copilots generating thousands of UIs daily—this verbosity becomes expensive fast.

OpenUI emerged from this observation with a controversial approach: what if we abandon JSON entirely and create a whitespace-sensitive DSL optimized for token efficiency? Instead of leveraging LLMs' native JSON capabilities, OpenUI teaches models a bespoke syntax through system prompts, trading ecosystem compatibility for a 52-67% reduction in tokens. It's a compiler-first approach to generative UI where the entire stack—from component definitions to streaming parsers—exists to support this custom language.

Technical Insight

The architecture revolves around three interconnected pieces: component library definitions, prompt generation, and the streaming parser. Unlike frameworks that generate JSON and parse it post-stream, OpenUI's parser operates on partial token streams, emitting renderable components mid-flight.

Component libraries are defined using Zod schemas that serve double duty—runtime validation and prompt generation. Here's what a component definition looks like:

import { z } from 'zod';
import { defineComponent } from '@openuidev/lang-core';

const ButtonComponent = defineComponent({
  name: 'Button',
  schema: z.object({
    variant: z.enum(['primary', 'secondary', 'danger']),
    children: z.string(),
    onClick: z.string().optional(),
  }),
  render: (props) => `<button class="btn-${props.variant}">${props.children}</button>`
});

When you register this component, OpenUI converts it into system prompt instructions that teach the LLM OpenUI Lang syntax. The generated prompt looks roughly like:

Button component syntax:
Button [variant] [children]
  variant: primary|secondary|danger
  children: string
  onClick: optional string

Example:
Button primary Submit
  onClick handleSubmit

The LLM then generates syntax like:

Form
  Button primary Submit
    onClick handleSubmit
  Button secondary Cancel

Notice the token savings: no quotes, no braces, no repeated keys. Indentation conveys hierarchy, position conveys property assignment. This is where the 60% reduction materializes—but it requires the LLM to perfectly nail indentation and positional semantics.

The streaming parser is where things get interesting architecturally. Instead of accumulating tokens and parsing complete payloads, OpenUI maintains a state machine that processes tokens incrementally:

class StreamingParser {
  private state: ParserState = { depth: 0, component: null, props: {} };
  private componentRegistry: ComponentRegistry;
  
  processToken(token: string): RenderableComponent | null {
    // Detect component start by checking registry
    if (this.componentRegistry.has(token)) {
      if (this.state.component) {
        // Emit previous component before starting new one
        const renderable = this.buildComponent(this.state);
        this.state = { depth: this.getIndentDepth(), component: token, props: {} };
        return renderable;
      }
      this.state.component = token;
      return null;
    }
    
    // Handle indentation changes
    if (token.startsWith('\n')) {
      const newDepth = this.getIndentDepth(token);
      if (newDepth < this.state.depth) {
        // Depth decrease means component complete
        return this.buildComponent(this.state);
      }
    }
    
    // Accumulate props
    this.state.props[this.inferPropKey()] = token;
    return null;
  }
}

This state machine approach enables sub-100ms first-paint because the moment the parser recognizes a complete component from partial tokens, it emits a renderable element. Traditional JSON streaming must wait for closing braces.

The framework adapters (React, Vue, Svelte) wrap this core parser with framework-specific rendering. The React implementation maintains a component tree that updates as new elements stream in:

export function useOpenUIStream(streamUrl: string) {
  const [components, setComponents] = useState<ReactElement[]>([]);
  const parserRef = useRef(new StreamingParser(componentRegistry));
  
  useEffect(() => {
    const stream = new EventSource(streamUrl);
    stream.onmessage = (event) => {
      const renderable = parserRef.current.processToken(event.data);
      if (renderable) {
        setComponents(prev => [...prev, renderToReact(renderable)]);
      }
    };
  }, [streamUrl]);
  
  return components;
}

The multi-framework strategy is genuinely architecture-agnostic. The lang-core package has zero React dependencies—it's pure TypeScript that emits an intermediate representation. Framework packages map this IR to their respective rendering primitives. Vue gets h() calls, Svelte gets component constructors, React gets createElement().

What makes this compelling for agent workflows is the LangChain integration. OpenUI exposes UI generation as a tool that agents can invoke:

import { OpenUITool } from '@openuidev/langchain';
import { ChatOpenAI } from 'langchain/chat_models/openai';

const tool = new OpenUITool({
  componentLibrary: myComponents,
  framework: 'react'
});

const agent = createReactAgent({
  llm: new ChatOpenAI(),
  tools: [tool]
});

// Agent can now generate UI as part of workflows
const result = await agent.invoke({
  input: "Show the user a form to collect shipping address"
});

The agent receives OpenUI generation as a callable function, treats UI as output, and the streaming parser handles incremental rendering. This positions OpenUI for orchestration scenarios where UI generation is one step in a multi-stage workflow.

Gotcha

The custom DSL is both OpenUI's superpower and its Achilles' heel. Smaller models struggle with the syntax—GPT-3.5 frequently generates malformed indentation that cascades into parser errors and blank screens. You're fighting against the LLM's training distribution, which has seen billions of JSON examples but zero OpenUI Lang examples. Every system prompt token spent teaching the syntax is a token not spent on domain knowledge.

The lack of server-side rendering is a production showstopper for many use cases. Because the streaming parser runs client-side, users see nothing until JavaScript executes and tokens start arriving. If your LLM takes 2 seconds to respond, that's 2 seconds of blank screen. There's no way to prerender a skeleton or static fallback because the component structure is unknowable until the LLM generates it. For public-facing applications, this destroys SEO and perceived performance. The framework only makes sense for authenticated dashboards or tools where JavaScript-required UX is acceptable.

Component library maintenance is pure toil. You're manually writing Zod schemas that mirror your actual React components, and there's no codegen or type-checking bridge between them. When you update a component's props, you must remember to update the schema, update the prompt template, and pray the LLM adapts. With 8,600 stars but minimal production case studies in the repo, this suggests the maintenance burden becomes prohibitive at scale.

Verdict

Use if: You're generating thousands of UIs per day and token costs are a measurable line item in your P&L. You control the component library completely (no third-party components). Your users are authenticated and JavaScript-required is acceptable. You have engineering capacity to maintain Zod schemas and debug LLM syntax errors. You're on GPT-4 or Claude Opus where syntax adherence is strong. Skip if: You're prototyping or have <500 generations/day—the token savings don't justify the complexity. You need SEO, server-side rendering, or fast time-to-first-paint. You're using open-source models below 70B parameters (syntax failure rates spike). You want ecosystem compatibility with standard React tooling. Vercel AI SDK already meets your needs—don't optimize prematurely.