Reproducible low-RAM benchmark (CPU-only, WikiText-2, llama.cpp build 10516 / b95502ba9).
Full data & method: b4ph/qwen3-4b-lowram-bench
| Quant | Size | PPL (WikiText-2) | Δ vs F16 | pp512 t/s | tg128 t/s | Verdict |
|---|
All numbers CPU-only (16 threads, AVX-512, -ngl 0, ctx 2048, seed 1).
PPL = perplexity (lower is better). pp512 = tokens/s processing a 512-token prompt.
tg128 = tokens/s generating 128 tokens. Machine-readable: results.tsv.
Base model: Qwen/Qwen3-4B (Apache-2.0). Community benchmark, not affiliated with Qwen or Hugging Face.