> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

Running 35B MoE Models on 8GB GPUs: How vvllm-ms Caches Experts in System RAM

[ View on GitHub ]

Running 35B MoE Models on 8GB GPUs: How vvllm-ms Caches Experts in System RAM

Hook

A Qwen 35B sparse model running on a laptop RTX 4060 sounds impossible—until you realize most MoE experts stay cold during math reasoning, making them perfect candidates for PCIe-backed caching.

Context

Mixture-of-Experts models like Qwen-MoE and DeepSeek-V2 deliver incredible performance by routing tokens through specialized expert networks, but they create a memory crisis for consumer hardware. A 35B-parameter sparse model might have 64 experts where only 2-8 activate per token, yet traditional inference engines like vLLM require all expert weights to live in VRAM simultaneously. This means a model with 8 active experts but 64 total experts needs the same memory as a 280B dense model—far beyond consumer GPU capacity.

The implicit assumption has always been that MoE inference requires enterprise cards with 40GB+ VRAM or multi-GPU setups with tensor parallelism. But what if expert selection patterns exhibit locality? If certain experts dominate during specific tasks (math experts for GSM8K, code experts for HumanEval), then cold experts are wasting precious VRAM. vvllm-ms exploits this insight by treating experts as cacheable units that can be paged between GPU memory and CPU RAM, similar to how operating systems page virtual memory to disk—except with microsecond PCIe transfers instead of millisecond disk seeks.

Technical Insight

The architecture introduces a paged expert cache layer that intercepts MoE router decisions before expert computation. When vLLM's scheduler selects experts for a token, the cache manager checks if those expert bundles (weights, biases, and critically, quantization scales for NVFP4/INT4 formats) reside in VRAM. Cache misses trigger an LRU eviction policy: the least-recently-used expert bundles get DMAed to pinned system RAM, freeing VRAM for the incoming hot experts. The key innovation is treating quantization metadata as part of the atomic cache unit—transferring FP16 weights without their INT4 scales would cause silent accuracy degradation.

Here's how you'd configure vvllm-ms for a Qwen-MoE model on an 8GB GPU:

from vllm import LLM, SamplingParams
import os

# Disable FlashInfer sampler to avoid kernel conflicts
os.environ['VLLM_USE_FLASHINFER_SAMPLER'] = '0'

llm = LLM(
    model="Qwen/Qwen1.5-MoE-A2.7B-NVFP4",
    trust_remote_code=True,
    max_model_len=4096,
    max_num_seqs=1,  # Single-sequence only
    tensor_parallel_size=1,  # No TP support
    gpu_memory_utilization=0.85,
    # Expert cache configuration
    moe_expert_cache_size_gb=6.0,  # 6GB VRAM for experts
    moe_backend='emulation',  # Router interception mode
    enforce_eager=True,  # No CUDA graph compilation
    download_dir="/path/to/checkpoints",
    # Requires complete checkpoint validation
    skip_tokenizer_init=False
)

sampling = SamplingParams(temperature=0.7, top_p=0.9, max_tokens=512)
outputs = llm.generate(["Calculate 347 * 892 step by step"], sampling)

The moe_backend='emulation' parameter is the tell: rather than modifying CUDA kernels, they're hooking the router's expert selection logic at the Python/C++ boundary. When the router returns expert indices, the cache manager intercepts that list, checks VRAM residency, and orchestrates transfers before passing control to the actual expert computation kernels. This design avoids forking vLLM's CUDA codebase, making it easier to track upstream changes, but it adds latency at the Python boundary that a kernel-level implementation could avoid.

The deterministic wave execution model is an opinionated choice for reliability. When a batch of tokens needs more experts than fit in the cache, the system doesn't attempt optimistic prefetching or partial batches—it blocks, performs all necessary transfers, then executes. This creates bimodal latency: cache hits deliver sub-millisecond expert access, while cache misses might take 10-50ms depending on PCIe generation and bundle size. For interactive single-user inference, this predictability beats the alternative of thrashing the cache with failed prefetch speculation.

Benchmark results on GSM8K (grade school math) demonstrate why this works for specific workloads. Math reasoning heavily reuses a small subset of experts—likely arithmetic and symbolic manipulation specialists—creating 80%+ cache hit rates. The system spends most inference time with all active experts already in VRAM, with occasional PCIe penalties when transitioning between problem types. Creative writing or diverse-domain QA would thrash this cache mercilessly, but structured reasoning tasks exhibit the locality this architecture exploits.

One subtle detail: the forced offline mode (download_dir required, no streaming downloads) exists because partial checkpoint downloads caused silent corruption during expert loading. When an expert bundle transfers from CPU to GPU but the weights file was incomplete, the original implementation would sometimes load garbage data into VRAM without failing loudly. By requiring complete checkpoint validation before engine startup, they trade convenience for correctness—you can't start inference until all expert bundles are verified on disk.

Gotcha

The single-GPU eager execution constraint is a fundamental scalability ceiling. tensor_parallel_size=1 means no tensor parallelism, and eager mode means no CUDA graph optimizations that vLLM normally uses for throughput. You cannot batch multiple requests (max_num_seqs=1 is hardcoded in practice), making this useless for production serving where amortizing compute across concurrent requests is critical. If your use case involves API endpoints serving multiple users, standard vLLM with a smaller dense model will deliver 10-100x better throughput.

Model compatibility is brittle and underdocumented. The examples show only Qwen NVFP4 and Gemma models working, suggesting the cache manager makes assumptions about expert topology that break on architectures with shared experts (like DeepSeek-V2's routed and shared expert split), grouped-query attention variants, or non-standard MoE layers. The LRU eviction policy is also naive—it has no awareness of upcoming expert needs based on attention patterns or prompt structure. A smarter system would prefetch experts likely to activate based on the current attention context, but implementing that would require model-specific profiling and prediction logic that doesn't exist here. You're stuck with reactive caching that always pays the PCIe tax on first access.

Verdict

Use if: You're running interactive experiments with Qwen or Gemma MoE models on 8-16GB consumer GPUs, working with structured reasoning tasks (math, code, logic puzzles) that exhibit high expert locality, and can tolerate variable latency spikes during expert transitions. This is perfect for researchers doing single-user model exploration where memory constraints matter more than throughput. Skip if: You need production serving with batching and predictable p99 latency, want multi-GPU support or tensor parallelism, or your workload involves creative generation and diverse-domain QA where expert usage patterns are chaotic. For those cases, standard vLLM with a properly-sized dense model (Llama 8B/13B on 16GB, Llama 70B on A100s with TP) delivers better reliability and performance. Also skip if you're working with MoE architectures beyond Qwen/Gemma—the lack of explicit compatibility documentation means you'll waste hours debugging silent failures on untested model families.