Qwen3-4B — pick your quant

Reproducible low-RAM benchmark (CPU-only, WikiText-2, llama.cpp build 10516 / b95502ba9). Full data & method: b4ph/qwen3-4b-lowram-bench

Size vs accuracy

Size vs speed (tok/s, generation)

Benchmarks (click a row to pick)

QuantSizePPL (WikiText-2) Δ vs F16pp512 t/stg128 t/sVerdict
💡 imatrix finding: at Q3_K_M, imatrix quantization buys back 0.84 PPL (15.66 → 14.82) for a one-time calibration pass. At Q4_K_M the gain is small (+0.05). Rule of thumb: run imatrix when you go Q3 or below.

Your pick:

File size
Perplexity (WikiText-2)
Quality cost vs full precision
Prompt processing
Generation speed
Recommended RAM
Download GGUF

All numbers CPU-only (16 threads, AVX-512, -ngl 0, ctx 2048, seed 1). PPL = perplexity (lower is better). pp512 = tokens/s processing a 512-token prompt. tg128 = tokens/s generating 128 tokens. Machine-readable: results.tsv. Base model: Qwen/Qwen3-4B (Apache-2.0). Community benchmark, not affiliated with Qwen or Hugging Face.