> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← Back to Articles

HyperQwen: 1,000 tok/s on a Single RTX 3090 Through Aggressive Recurrent-State Speculation

[ View on GitHub ]

HyperQwen: 1,000 tok/s on a Single RTX 3090 Through Aggressive Recurrent-State Speculation

Hook

A production-grade LLM serving stack that hits 382 tokens per second when editing documents—by letting the model draft tokens from your own prompt instead of predicting them.

Context

Serving large language models on consumer hardware has always been a game of memory Tetris. The standard advice—'just rent an A100' or 'quantize to 4-bit and pray'—ignores a massive market: developers and small teams running inference workloads on RTX 3090s and 4090s who need production throughput without cloud bills. Stock vLLM gets you continuous batching and PagedAttention, but on a single 24GB card with Qwen 3.8-27B, you're looking at 30-40 tokens per second single-user, maybe 200-300 aggregate with heavy batching.

The problem compounds with newer recurrent-state architectures like Mamba and its variants. Qwen 3.8-27B uses DeltaNet, which avoids the quadratic KV cache growth of transformers—except existing serving frameworks treat recurrent state as an afterthought. Speculation strategies designed for transformer attention don't account for how recurrent state scales with verify block width. Prefix caching implementations checkpoint attention KV but discard the recurrent state, forcing expensive recomputation on every conversation turn. HyperQwen is a vLLM 0.28.0 patch series that rewrites the serving economics for this exact model-hardware pair: Qwen 3.8-27B on RTX 3090, squeezing out 3-8× throughput by treating memory pressure as the primary constraint and making every architectural decision trade precision or complexity for VRAM headroom.

Technical Insight

DFlash2 Speculation

Batch Mode Pipeline

Mode Selection

Request Entry

Multi-user

Single-user

Memory Optimization

INT4 Requantized

LM Heads

KVarN Cache

4/2-bit 240k ctx

Calibrated Vocab

8K vs 152K tokens

User Request

Scheduler

Batch Mode

Speculation Mode

INT8 Tensor GEMMs

FP16 DeltaNet State

~1000 tok/s

64 concurrent

Draft Model

7 model tokens

Context Draft

7 prompt tokens

Split-KV Verify

15-token window

Verification

Output Tokens

System architecture — auto-generated

HyperQwen's core insight is that recurrent-state models fundamentally change speculation economics. In transformer speculation, the bottleneck is attention verification over draft tokens—but KV cache grows with context length, not draft width. With DeltaNet, recurrent state grows with verify block width. Speculating 15 tokens instead of 7 means each resident request consumes 1.66 GiB of state versus 0.88 GiB. This makes wide speculation a single-user feature, not a batching win.

The architecture splits into two serving modes. Batch mode prioritizes concurrent throughput: int8 tensor-core GEMMs for matrix operations (trading fp16 precision for 2× memory reduction), fp16 DeltaNet state, and calibrated draft vocabularies. This hits ~1,000 tok/s at 64 concurrent requests. Single-user mode implements DFlash2 block-draft speculation: the drafter proposes 7 tokens per forward pass, but positions 7-14 draft from the request's own context instead of model predictions. Here's the configuration knob:

# In vllm/core/scheduler.py - DFlash2 speculation config
class SpeculationConfig:
    draft_tokens: int = 7  # Model-predicted tokens
    context_tokens: int = 7  # DFLASH_TOKENS: draft from prompt
    verify_width: int = 15  # Total speculation window
    
    # Draft vocab constrained to empirically observed tokens
    # Calibrated on target domain (code/prose/math)
    draft_vocab_size: int = 8192  # vs 152k full vocab

The context-aware speculation is devastatingly effective for document reproduction—when your workload involves editing, templating, or code refactoring where the output heavily overlaps the input. Benchmarks show 382 tok/s versus 132 tok/s without it. But on prose generation, context drafts contribute only 0.65% of accepted tokens, making DFLASH_TOKENS a per-workload parameter, not a universal win. The system doesn't auto-tune this; you hardcode it based on your use case.

Prefix caching demonstrates the recurrent-state preservation advantage. Standard vLLM caches attention KV but discards DeltaNet state between turns. HyperQwen checkpoints the recurrent state alongside attention context:

# Simplified recurrent-state caching logic
class RecurrentStateCache:
    def checkpoint(self, request_id, state_tensor):
        # state_tensor: [batch, hidden_dim] DeltaNet state
        self.cache[request_id] = {
            'state': state_tensor.clone(),
            'position': self.current_position,
            'hash': self._hash_prefix(request_id)
        }
    
    def restore(self, request_id, new_prompt):
        cached = self.cache.get(request_id)
        if cached and self._hash_prefix(new_prompt) == cached['hash']:
            # Resume from checkpointed state instead of recomputing
            return cached['state'], cached['position']
        return None, 0

This cuts second-turn latency from 22.4 seconds to 0.56 seconds—critical for agentic workflows where the model makes multiple calls with shared context. The cache key includes a prefix hash, so conversations with identical opening turns share state across requests.

Memory pressure optimization is relentless. The lm_head (projection from hidden states to vocabulary logits) normally uses fp16, consuming ~800 MB for Qwen's 152k vocab. HyperQwen requantizes it to calibrated int4: measure activation ranges on a representative dataset, compute per-channel scales, quantize weights to 4-bit. The requantization script (scripts/requant_lm_head.py) takes a GGUF checkpoint and produces a quantized safetensors file. This saves ~600 MB with negligible quality degradation because the top-k sampling only needs relative logit ordering, not absolute precision.

For long context (150k-262k tokens), KVarN integration applies 4/2-bit lossy compression to the KV cache. The key architectural question: does lossy KV compression break speculation? If the draft model sees quantized KV and the verify model sees different quantized KV (due to recomputation), will acceptance rates tank? HyperQwen's answer: recurrent state captures enough information that quantization noise in attention doesn't cascade into draft rejections. GSM8K accuracy stays at 95-96.5% across quantization settings, though this only validates math reasoning, not retrieval or multi-turn coherence.

The forced migration to vLLM's V2 model runner exposes infrastructure fragility. DFlash2 requires UVA (Unified Virtual Addressing) buffer allocation before weight loading—a requirement the V1 runner doesn't satisfy. This surfaces WSL2 paravirtualization overhead: the Linux kernel's view of GPU memory doesn't match WDDM's actual allocations, causing 20% throughput loss and 12 GB usable VRAM instead of 24 GB. Running bare-metal Linux recovers the performance, but native Windows CUDA never reaches parity. This isn't a tuning issue—it's a hard deployment constraint that the README buries in troubleshooting notes.

Gotcha

HyperQwen is welded to Qwen 3.8-27B and RTX 3090. The requantization scripts hardcode layer names and tensor shapes. The calibration dataset for draft vocabularies and lm_head quantization is unspecified—'representative data' could mean instruct samples, code, or math problems, and your quality will silently degrade if the calibration domain mismatches production. Adapting this stack to Qwen 3.8-72B or a 40GB A100 requires rewriting memory budgets, re-running calibration, and tuning speculation parameters. There's no abstraction layer; it's a reference implementation demonstrating what's possible, not a framework.

Quality evaluation is dangerously narrow. GSM8K exact-match (n=200 math problems) and one 200k-token needle test don't validate that 4/2-bit KV quantization preserves multi-turn coherence, retrieval across diverse contexts, or specialized domain performance. Production use demands contamination-free evals across instruction-following, code generation, long-context reasoning, and safety. The recurrent-state caching has no analysis of state drift over long conversations—does the checkpointed state degrade after 50 turns? The benchmarks don't say. Concurrent throughput claims lack reproducible protocols: request arrival patterns, output length distributions, cold-start versus steady-state measurements. 'Aggregate decode at 1,035 tok/s' could mean best-case batching with pre-warmed caches, and until recently, prefill numbers were contaminated by prefix cache hits. You can't bet production SLAs on these numbers without running your own benchmarks.

Verdict

Use if: You're running Qwen 3.8-27B on RTX 3090/4090 for production workloads and can tolerate tight hardware coupling. Your use case involves document editing, code refactoring, or template expansion where context-aware speculation wins big. You have engineering capacity to maintain a vLLM fork and re-run calibration when Qwen releases new checkpoints. You need 3-8× throughput over stock vLLM and cloud costs make GPU rental unattractive. Skip if: You need multi-model flexibility or plan to support multiple architectures—this is a one-model, one-GPU solution with no portability path. Your workloads require quality guarantees beyond toy benchmarks; the evaluation is too shallow for regulated domains or safety-critical applications. You're on WSL2 and can't migrate to bare-metal Linux; the 20% throughput tax and VRAM loss kill the value proposition. You expect framework-level abstractions—HyperQwen is a patch series and requantization cookbook, not a drop-in serving solution. For production fleets, wait for vLLM to upstream the wins or budget engineering time to maintain the fork yourself.