> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

FreeToken: Bandwidth-Adaptive Expert Scheduling for MoE Models on Consumer GPUs

[ View on GitHub ]

FreeToken: Bandwidth-Adaptive Expert Scheduling for MoE Models on Consumer GPUs

Hook

Running DeepSeek-V3's 671B parameters on a single RTX 4090 sounds impossible—until you realize that MoE models only activate 2-8% of their weights per token, and PCIe 4.0 can stream 32GB/s if you're smart about it.

Context

Mixture-of-Experts models changed the economics of frontier AI. Instead of activating all 405B parameters like Llama 3.1, DeepSeek-V3 routes each token through just 37B of its 671B total parameters, achieving GPT-4-class performance at a fraction of the compute cost. But this architectural win creates a deployment nightmare: expert routing is unpredictable, memory requirements spike based on which experts fire, and traditional inference engines either keep all experts in VRAM (impossible on consumer hardware) or use naive CPU offloading that kills latency.

Existing solutions pick bad tradeoffs. vLLM and TensorRT-LLM assume datacenter GPUs with 80GB+ VRAM. llama.cpp and ExLlamaV2 optimize for dense models and treat MoE as an afterthought. DeepSpeed-Inference has flexible offloading, but its static partitioning can't adapt when PCIe bandwidth fluctuates or when a coding session suddenly needs 10x more KV cache for a giant file. The prosumer hardware sitting in ML engineers' workstations—RTX 4090s with 24GB VRAM and PCIe 4.0 x16 slots—has enough raw bandwidth to stream experts on-demand, but no inference engine exploits it intelligently.

Technical Insight

FreeToken's core innovation is the q* bandwidth-adaptive policy, a runtime orchestrator that profiles actual PCIe throughput every few hundred milliseconds and dynamically decides expert placement. Unlike static partitioning schemes that commit to 'these experts live on GPU, those on CPU' at engine startup, q* measures real-world transfer speeds and routes computation accordingly. If your system is hitting 28GB/s PCIe bandwidth, it aggressively streams cold experts from host RAM. If background processes saturate the bus and bandwidth drops to 12GB/s, it shifts toward GPU-resident caching.

The architecture implements a three-tier memory hierarchy that looks deceptively simple but handles the hardest edge cases. Tier 1 is GPU VRAM for hot experts—the 5-10% of experts that fire most frequently get pinned and never evicted. Tier 2 is host RAM with a global LRU cache shared across all layers. This is critical: MoE routing exhibits strong temporal locality. If expert_42 fires in layer 10, there's a 60-80% chance it fires in layer 11. Traditional per-layer caching misses this cross-layer pattern entirely. Tier 3 is double-buffered streaming where the scheduler issues async PCIe transfers for the next layer's predicted experts while the current layer computes, hiding transfer latency through overlap.

Here's what the execution flow looks like for a single forward pass:

# Simplified FreeToken scheduling logic
class MoEScheduler:
    def __init__(self, model, vram_budget):
        self.expert_cache = LRUCache(capacity=vram_budget)
        self.bandwidth_oracle = PCIeBandwidthProfiler()
        self.prefetch_buffer = DoubleBuffer(size=2 * expert_size)
        
    async def forward_layer(self, layer_idx, hidden_states, routing_weights):
        # Routing決定 which experts to activate
        active_experts = torch.topk(routing_weights, k=2, dim=-1)
        
        # q* policy: decide CPU vs GPU execution
        current_bw = self.bandwidth_oracle.measure()  # Real-time PCIe speed
        transfer_cost = expert_size / current_bw
        compute_cost = self.estimate_gemm_latency(hidden_states)
        
        expert_outputs = []
        for expert_idx in active_experts:
            if expert_idx in self.expert_cache:
                # Tier 1: GPU-resident, direct execution
                expert_outputs.append(
                    self.expert_cache[expert_idx](hidden_states)
                )
            elif transfer_cost < compute_cost * 0.3:
                # Tier 2: Stream from CPU, bandwidth is good
                expert_weights = await self.prefetch_buffer.get(expert_idx)
                expert_outputs.append(self.execute_on_gpu(expert_weights, hidden_states))
            else:
                # Tier 3: Execute on CPU, PCIe is saturated
                expert_outputs.append(
                    await self.execute_on_cpu(expert_idx, hidden_states.cpu())
                )
        
        # Prefetch next layer's likely experts based on routing history
        self.prefetch_buffer.async_load(
            self.predict_next_experts(layer_idx + 1)
        )
        
        return self.merge_expert_outputs(expert_outputs, routing_weights)

The bandwidth oracle is where theory meets hardware reality. PCIe 4.0 x16 advertises 32GB/s, but real sustained bandwidth depends on chipset configuration, CPU PCIe lane allocation, whether your NVMe drives are stealing lanes, and even motherboard trace quality. FreeToken profiles this continuously by timing actual cudaMemcpyAsync calls and maintaining a rolling average with outlier rejection. When bandwidth drops—say, you start a file transfer or a Docker build—the scheduler detects it within 200-300ms and shifts toward more GPU-resident caching.

Semantic anchor checkpointing solves a different problem: agentic workflows waste compute. When an LLM coding assistant generates 500 tokens of Python, realizes it made a mistake, and wants to retry, traditional KV caches offer no help. You either recompute the entire context or keep redundant cache states in memory. FreeToken's semantic anchors create content-addressed snapshots at logical boundaries—end of system prompt, after each tool call, before each code generation block. The checkpointing system hashes the concatenated KV cache and model state to create a 256-bit anchor ID:

def create_semantic_anchor(kv_cache, expert_state, prompt_boundary):
    # Content-addressable checkpoint at logical boundary
    checkpoint = {
        'kv_cache': kv_cache.detach().clone(),
        'expert_lru_state': expert_state.snapshot(),
        'position': prompt_boundary,
        'hash': sha256(kv_cache.tobytes() + expert_state.serialize())
    }
    return checkpoint['hash'], checkpoint

# Rollback is O(1) lookup, not O(n) recomputation
def rollback_to_anchor(anchor_hash):
    checkpoint = anchor_store[anchor_hash]
    model.load_kv_cache(checkpoint['kv_cache'])
    scheduler.restore_expert_state(checkpoint['expert_lru_state'])
    return checkpoint['position']  # Resume generation from here

When the model wants to retry a tool call, it rolls back to the pre-call anchor, avoiding full context recomputation. This is transformative for coding assistants that iterate: instead of paying for 2000 tokens of recomputation, you pay for a hash lookup and a memory copy.

The FTW (FreeToken Weight) format deserves attention because it's optimized for the access pattern of streaming MoE inference. Unlike GGUF's monolithic file structure or SafeTensors' tensor-per-file organization, FTW stores each expert as an independent compressed block with its own metadata header. This enables expert-granular decompression—you only decompress the 2-4 experts needed for the current token, not the entire layer. The format supports mixed quantization per expert: you can store critical experts in FP8 for accuracy and rarely-used experts in MXFP4 to save bandwidth. This granularity is impossible with model-level quantization schemes.

Gotcha

The NVIDIA-only limitation isn't just about CUDA versus ROCm—it's architectural. The q* policy depends on NVLink/PCIe performance characteristics, cudaMemcpyAsync behavior, and cuBLAS kernel timing that don't translate to AMD or Intel GPUs. The bandwidth oracle's profiling assumptions break on discrete AMD cards where memory transfers have different latency profiles. Porting this to ROCm would require re-profiling and retuning the entire scheduling heuristic.

Semantic anchor checkpointing has a subtle correctness problem that the documentation doesn't address: hash collisions or near-duplicate contexts. If two different conversation states hash to the same 256-bit anchor (astronomically unlikely but not impossible), or if a coding assistant generates nearly identical code blocks with slight variations, the rollback mechanism could restore the wrong state. There's no validation that the restored checkpoint actually matches the intended semantic boundary. For production agentic systems that make decisions based on subtle context differences, this is a latent bug waiting to surface. The codebase would benefit from checkpoint validation—store a lightweight semantic signature (first/last 512 bytes of KV cache) alongside the hash and verify on rollback.

Verdict

Use if: You're building coding assistants or tool-calling agents that need to run 200B+ parameter MoE models (DeepSeek-V3, Qwen-MoE, Mixtral 8x22B) on prosumer NVIDIA hardware (RTX 4090, 4080, or multi-GPU 3090 setups), you have interactive latency requirements where users expect <2 second response times, or you're running agentic workflows where the LLM frequently backtracks and retries tool calls. FreeToken's bandwidth-adaptive scheduling and semantic checkpointing directly solve these use cases in ways that vLLM and llama.cpp don't. Skip if: You're deploying dense models like Llama or Mistral where the expert-scheduling machinery is pure overhead, you need production-grade stability and monitoring (this project is too young for mission-critical serving), you're running on AMD/Intel/Apple hardware where the CUDA-specific optimizations don't apply, or you have datacenter GPUs with 80GB+ VRAM where keeping all experts resident is trivial. Also skip if you need multi-user batched serving—FreeToken optimizes for single-session interactive latency, not throughput-oriented batch inference.