Qwen3.8-27B on a Single RTX 3090: How Calibrated Quantization and Speculative Decoding Hit 1,000 tok/s
Hook
A 27-billion parameter model hitting 1,000 tokens per second on a $1,500 consumer GPU shouldn't be possible—but calibrated int4 quantization of just the output projection changes the math entirely.
Context
Running frontier-scale language models locally has always been a game of compromises. You either quantize aggressively and watch quality degrade, or you accept glacial inference speeds that make interactive applications impossible. The standard playbook—naive int8 quantization, static KV caching, sequential token generation—leaves massive performance on the table because it treats all model components equally. A linear layer processing hidden states has different numerical sensitivity than the final output projection. The KV cache storing attention history has different access patterns than recurrent state in hybrid architectures.
Qwen3.8-27B compounds these challenges. Its DeltaNet recurrent layers (a hybrid attention-RNN architecture) add state management complexity that vanilla vLLM doesn't handle. The model's 27 billion parameters barely fit in a 3090's 24GB VRAM even when quantized, leaving no room for aggressive batching or long context windows. Most deployment guides for models this size assume you're renting A100s or settling for 20-30 tok/s on consumer hardware. This repository demolishes those assumptions by treating quantization and speculative decoding not as blunt instruments but as surgical optimizations applied exactly where they matter most.
Technical Insight
The performance breakthrough comes from heterogeneous quantization—different precisions for different components based on measured sensitivity. While the recurrent DeltaNet state stays in fp16 to preserve quality across long sequences, all linear layers drop to int8 using tensor-core GEMMs. But the real innovation is the calibrated int4 lm_head. In speculative decoding, the verification step dominates cost because you're running the full model's output projection on every draft token to accept or reject it. The standard approach keeps this projection in int8 or fp16, which means 75% of your verification bandwidth goes to a single matrix multiply.
The calibration process uses activation statistics from real prompts rather than naive zero-point quantization:
# From prepare/requant_lm_head.py (simplified)
import torch
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.8-27B")
lm_head = model.lm_head.weight # [vocab_size, hidden_dim]
# Collect activations from calibration dataset
calib_acts = []
for prompt in calibration_prompts:
hidden_states = model.forward(prompt, output_hidden_states=True)
calib_acts.append(hidden_states[-1]) # Last layer activations
acts_concat = torch.cat(calib_acts, dim=0)
# Compute per-channel scales from actual activation distribution
scales = acts_concat.abs().max(dim=0)[0] / 7.0 # int4 range is [-7, 7]
# Quantize using calibrated scales, not naive min/max
quantized_weight = torch.clamp(
torch.round(lm_head / scales.unsqueeze(0)),
-7, 7
).to(torch.int8)
This dataset-aware quantization preserves the output distribution's tail behavior—the rare but critical logits that steer generation quality. The result: speculative decoding acceptance rates stay above 90% even with a 4-bit verification head, cutting memory bandwidth by 75% without the quality collapse you'd see from naive quantization.
The split-KV verify attention is equally clever. Standard speculative decoding recomputes attention for all draft tokens during verification because the drafter and verifier might have different KV caches. This repo exploits the fact that accepted tokens produce identical KV pairs in both models—so it caches them once and reuses across verification steps. The implementation patches vLLM's attention kernel to split the KV cache into 'verified' and 'speculative' regions:
# Conceptual patch to vLLM's attention (actual patch is in CUDA)
class SplitKVAttention:
def __init__(self):
self.verified_kv = [] # Accepted tokens
self.draft_kv = [] # Speculative tokens
def verify_step(self, draft_tokens, model):
# Only compute attention for new/rejected tokens
verify_outputs = model(draft_tokens,
kv_cache=self.verified_kv)
accepted = self.compare_logits(verify_outputs, draft_tokens)
# Move accepted KV pairs from draft to verified
self.verified_kv.extend(self.draft_kv[:accepted])
self.draft_kv = [] # Clear rejected drafts
return accepted
The draft-from-context optimization (DFLASH_TOKENS=15) takes this further by recognizing that LLMs often quote their input verbatim. During verification, positions 7-14 are filled by copying directly from the prompt rather than running the drafter. On document summarization workloads where the model reproduces passages, this hits 382 tok/s versus 260 tok/s without it—a 47% gain. The trade-off is memory: each request now needs double the recurrent-state pages because you're tracking both the draft context and the original prompt. The README explicitly advises enabling this only for quote-heavy workloads, showing real operational awareness of the concurrency cost.
The KVarN integration extends context to 268k tokens via lossy 4/2-bit KV cache compression. Instead of storing every attention key and value at fp16 (2 bytes per element), KVarN clusters similar KV pairs and stores cluster indices plus residuals. For a 150k token context, this compresses ~18GB of cache down to ~4GB. The compression is transparent to the model—the attention kernel decompresses on-the-fly during each forward pass. The risk is quality degradation on long-context tasks, which the repo validates with a 200k needle-in-haystack test and GSM8K scores that stay within 95.0-96.5% across all configurations.
All of this runs in a Docker container where 19 vLLM patches are applied at build time and verified by verify.sh before the image is tagged. First boot downloads the base Qwen3.8-27B checkpoint and runs the requantization scripts automatically, storing the calibrated int4/int8 weights in a named volume. You choose batch or single-user mode via Docker Compose profiles—they're mutually exclusive because the memory allocation strategies conflict.
Gotcha
This is a monolithic optimization for one model on one runtime. The 19 vLLM patches target commit a3f8b12 specifically, and the quantization scripts hardcode Qwen's architecture (layer names, hidden dimensions, DeltaNet recurrent structure). Porting this to Llama 3.3 or Mistral Large would require re-deriving the calibration statistics for a completely different activation distribution, rewriting the recurrent state logic (Llama doesn't have DeltaNet), and debugging new CUDA graph capture issues because each architecture has different residual connection patterns. There's no abstraction layer—you're modifying vLLM internals directly.
The WSL2 support is functional but fragile. You lose ~20% throughput versus bare metal Linux, usable VRAM drops by 500MB (which matters when you're already at 23.5GB/24GB utilization), and exceeding memory doesn't fail fast with OOM—it triggers silent WDDM paging that causes 2-6x slowdowns. The environment variable VLLM_WSL2_ENABLE_PIN_MEMORY=1 is mandatory or the V2 runner crashes with 'UVA is not available', but that flag itself increases memory pressure. One contributor measured native Windows as even slower than WSL2 (66.0 vs 76.2 tok/s), so there's no escape hatch. If you're deploying on WSL2, you need to know your exact memory budget down to the hundred-megabyte level.
KVarN's lossy KV cache compression has thin validation. One 200k needle test and a GSM8K benchmark that 'sits inside the band' isn't enough coverage for a compression scheme storing 268k tokens of context. The fact that verbatim reproduction became 'bit-identical across runs' only after a PIECEWISE fix suggests there were correctness bugs that took specific workloads to expose. Lossy compression of recurrent state is fundamentally risky—you're trading context fidelity for capacity—and the test matrix doesn't include adversarial inputs or tasks that require precise long-range reasoning.
Verdict
Use if: You're deploying Qwen3.8-27B specifically, need either high single-user interactivity (100+ tok/s) or batch throughput (1,000+ tok/s at 64 concurrent), have bare-metal Linux or well-understood cloud GPU access, and your workload cleanly fits one profile (you're not trying to dynamically balance interactive and batch traffic). This is the blueprint for calibrated quantization and surgical speculative decoding—the patches and benchmarks teach you what to optimize even if you never run this exact stack. Skip if: You need multi-model support, plan to run on WSL2 in production (the memory footguns are documented but not solved), want dynamic workload balancing between single-user and batch modes, or can't commit to maintaining 19 custom patches against vLLM. The performance is exceptional but the generality is zero—this solves one model at one scale brilliantly and nothing else.