> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

Running a 2.78-Trillion-Parameter Model on a Laptop: Inside WASTE's NVMe-Streaming Architecture

[ View on GitHub ]

Running a 2.78-Trillion-Parameter Model on a Laptop: Inside WASTE's NVMe-Streaming Architecture

Hook

A MacBook Pro with 64GB RAM and a 1TB SSD can now run a model 40x larger than what fits in memory—not through clever compression, but by treating your NVMe drive as extraordinarily slow VRAM, streaming 17GB of expert weights per generated token.

Context

The race to run frontier language models locally hit a wall when Mixtral and DeepSeek introduced massive mixture-of-experts architectures. While MoE models activate only a fraction of their parameters per token (making them theoretically tractable), the inactive experts still need to live somewhere. Kimi's K3 model pushes this to the extreme: 2.78 trillion parameters where even the always-active 'trunk' consumes 27GB of RAM. The traditional playbook—aggressive quantization, model pruning, cloud offloading—doesn't work when you specifically need the full parameter count for quality or when cloud deployment is off the table due to privacy requirements or cost constraints.

WASTE (Weight Activation Streaming Engine) attacks this problem from first principles: if experts are selected dynamically and most stay dormant for any given token, why keep them in RAM at all? Instead of treating memory capacity as a hard constraint to work around, WASTE rebuilds the inference engine to stream expert weights from NVMe on-demand, overlapping I/O with computation through speculative scheduling. This isn't a hack bolted onto existing frameworks—it's a ground-up C implementation with no dependencies on BLAS, Python runtimes, or GPU libraries, designed around the assumption that PCIe bandwidth to your SSD, not FLOPs, is the bottleneck.

Technical Insight

The core architectural insight is the lookahead router, which treats expert prediction as a purely speculative scheduling hint. Before each layer processes a token, a lightweight predictor examines the hidden states and issues async reads for the experts it thinks will be needed. Meanwhile, the actual router—using the model's learned routing logic—makes the real selection. If the predictor guessed right, the expert weights arrive just as the compute unit needs them. If it guessed wrong, the correct expert gets read synchronously and throughput drops, but outputs remain bit-identical to a full-memory implementation. This is fundamentally different from speculative decoding, where wrong guesses change outputs: here, speculation only affects timing.

The container format is purpose-built for this access pattern. Each expert is laid out as a single aligned read to minimize seek overhead—you can't memory-map this like a standard model file. The on-disk structure looks like this in pseudocode:

// Simplified container layout
struct WasteContainer {
  Header header;              // Model metadata, expert count
  TrunkWeights trunk;         // Always-resident 48B parameters (27GB @ 4/8-bit)
  ExpertIndex index[N];       // Offset + size for each expert
  Expert experts[N];          // Aligned to 4KB boundaries
};

// Inference loop with async prefetch
void process_token(Token tok) {
  // Lookahead router predicts next layer's needs
  int predicted_experts[8];
  predict_experts(tok.hidden_state, predicted_experts);
  
  // Issue async reads (non-blocking)
  for (int i = 0; i < 8; i++) {
    async_read_expert(predicted_experts[i]);
  }
  
  // Actual routing decision
  int actual_experts[8];
  route_experts(tok.hidden_state, actual_experts);  // Deterministic
  
  // Wait for reads to complete, fall back to sync if mispredicted
  for (int i = 0; i < 8; i++) {
    ExpertWeights* weights = wait_for_expert(actual_experts[i]);
    apply_expert(tok.hidden_state, weights);
  }
}

The heterogeneous quantization strategy reflects actual bottleneck analysis. Experts use aggressive 3-bit residual vector quantization because they're bandwidth-bound anyway—streaming 17GB per token means you're waiting on NVMe throughput, not compute. Squeezing experts to 3-bit halves the I/O cost with minimal quality degradation (0.037 KL divergence in measurements). Meanwhile, the always-resident trunk uses gentler 4/8-bit quantization because it's compute-bound and lives in L2 cache. This is the inverse of typical quantization strategies that apply uniform precision across all weights.

Unused RAM becomes a bounded LRU cache for recently-accessed experts. The performance table in the repo reveals a counterintuitive failure mode: the 27GB working set with 8GB cache hits 4.81 tokens/second, but expanding to 128GB cache (more than the entire expert footprint) collapses throughput to 0.58 tok/s. This isn't a typo—it's the OS paging 101GB of cached experts to swap, thrashing the same NVMe drive you're trying to stream from. The optimal configuration allocates just enough cache to cover frequently-reused experts without triggering page-out, demonstrating that working set optimization matters more than raw capacity.

The vision processing reveals the real cost structure for multimodal inference. The vision tower processes an image in 15.7 seconds, but each of the 1,024 image position embeddings then costs 2.8 seconds to run through the language model—2,867 seconds total for a single image. This isn't a bug; it's the fundamental cost of attention over visual tokens when you're streaming weights at 0.6 tok/s. Multimodal isn't just 'bolt CLIP onto an LLM' when your architecture is I/O-bound.

Gotcha

The 'beyond available RAM' claim is technically true but practically misleading. You need 64GB of RAM to run K3 because the 48-billion-parameter trunk consumes 27.28GB resident, plus overhead for the KV cache and OS. You're not running a 2.78T model in 8GB—you're running a 48B always-resident trunk that happens to access 2.73T of streamed experts. The streaming architecture only helps with the expert weights, and those experts still require 1TB of internal NVMe storage. The performance measurements show external USB drives delivering 0.94 GB/s versus 12.78 GB/s for internal SSDs—a 13x gap that makes this completely unusable on external storage. If you're on a laptop with 512GB internal storage, this won't work.

The CPU-only inference path is a deliberate architectural choice that caps competitive throughput. There's no GPU acceleration in the critical path despite WASTE being written in 2025, because the bottleneck is PCIe bandwidth (streaming 17GB per token), not FLOPs. Adding GPU compute doesn't help when you're waiting on NVMe reads. This makes sense for the target workload, but it means llama.cpp running a quantized 70B model entirely in RAM will deliver 10-50x better throughput on the same hardware. WASTE's 0.6 tok/s is only acceptable if you specifically need K3's full parameter count and no smaller model will suffice. The BACKENDS.md file admits Metal and CUDA support are 'to be explored,' which suggests GPU offloading won't arrive soon. The format and API are explicitly unstable—the README warns that both are subject to breaking changes. This is a research artifact with honest documentation, not a production-ready library.

Verdict

Use WASTE if you need to run the full Kimi K3 model locally for research or privacy-critical applications, you have 64GB RAM and 1TB internal NVMe storage available, and you're willing to accept 0.6 tok/s throughput to get access to 2.78 trillion parameters. The lookahead router and heterogeneous quantization techniques are genuinely clever and worth studying even if you never deploy WASTE itself. Skip it if you want practical inference speed (llama.cpp with Q4 quantization delivers 10-30 tok/s for 70B models on the same hardware), if you're deploying to production (unstable API, rapid iteration, no GPU acceleration), if you're on a typical laptop without high-end specs, or if a smaller model would actually meet your needs. The real contribution here is proof that streaming MoE inference from NVMe is viable at all—but viability and practicality are different thresholds.