Measurements
Measured values for throughput, memory requirement, MTP parameters and multi-user operation.
All measurements on one machine: Ryzen AI MAX+ 395, Radeon 8060S (gfx1151), 128 GB LPDDR5X, Fedora 44, ROCm/HIP 7.1,
Samsung 990 PRO NVMe. Governor “powersave”, tuned “balanced” (not optimized). Noise between runs about 3%.
The raw data are in the project under bench/results/.
llama-bench (without MTP, HIP, stock fork)
| Quant | KV | ubatch | pp512 | tg128 |
|---|---|---|---|---|
| UD-Q4_K_XL | f16 | 512 | 412 t/s | 21.0 t/s |
| UD-Q4_K_XL | q8_0 | 512 | 420 t/s | 20.9 t/s |
| UD-Q4_K_XL | f16 | 2048 | 422 t/s | 21.3 t/s |
| UD-Q2_K_XL | f16 | 512 | 395 t/s | 24.3 t/s |
At 512 tokens, KV type and ubatch change nothing measurable. Decoding is bandwidth-limited: per token about 5 GB are read (10 of 512 experts plus the dense parts).
Memory requirement and decode with MTP (32k context, q8_0, dzannotti head)
| Engine | Quant | MTP | Load time | Decode | Acceptance | Memory requirement |
|---|---|---|---|---|---|---|
| Stock (auto) | Q2_K_XL | – | 20 s | 24.6 | – | 75.7 GB |
| Stock (auto) | Q2_K_XL | n3/p0.75 | 22 s | 34.8 | 86% | 78.8 GB |
| Stock (auto) | IQ3_XXS | n3/p0.75 | 23 s | 35.1 | 88% | 81.5 GB |
| Stock (auto) | IQ4_XS | n3/p0.75 | 26 s | 34.2 | 87% | 92.7 GB |
| Stock (auto) | Q4_K_XL | n3/p0.75 | – | OOM | – | over 107 GB |
| EngramHalo (none) | IQ3_XXS | – | 17 s | 23.1 | – | 52.8 GB |
| EngramHalo (none) | IQ3_XXS | n4/p0.75+ngram | 16 s | 34.4 | 79% | 57.4 GB |
| EngramHalo (none) | IQ4_XS | n4/p0.75+ngram | 20 s | 36.5 | 84% | 68.7 GB |
| EngramHalo (none) | Q4_K_XL | n4/p0.75+ngram | 28 s | 35.3 | 80% | 84.7 GB |
| EngramHalo (mmap) | Q4_K_XL | n4/p0.75+ngram | 39 s | 35.5 | 80% | 83.4 GB |
MTP fine-tuning (EngramHalo, Q4_K_XL, 3 prompts of 400 tokens each, reasoning medium)
| Configuration | Ø | Code | Prose | Reasoning | Acceptance |
|---|---|---|---|---|---|
| without MTP | 20.8 | 20.8 | 20.8 | 20.9 | – |
| n2 / p0.75 | 30.9 | 30.7 | 28.2 | 33.8 | 85% |
| n3 / p0.75 | 33.9 | 33.1 | 29.7 | 38.9 | 83% |
| n4 / p0.75 + ngram-mod (preset) | 34.7 | 32.1 | 31.2 | 41.0 | 80% |
| n4 / p0.0 | 36.0 | 32.5 | 31.1 | 44.5 | 55% |
| n6 / p0.75 + ngram-mod | 35.7 | 34.9 | 30.9 | 41.3 | 76% |
| n3 / p0.75, temp 0.6 | 35.3 | 36.1 | 30.4 | 39.5 | 86% |
| n3 / p0.75, temp 0.0 | 35.8 | 36.2 | 30.8 | 40.3 | 87% |
All values in tokens per second, temperature 1.0 unless stated otherwise. Reasoning outputs benefit the most.
Draft head and context depth (UD-IQ4_XS, one slot, temp 1.0, 600 output tokens)
| Draft head | Size | ~0 context | 4k | 8k | 30k | 64k |
|---|---|---|---|---|---|---|
| dzannotti Q4_K_M | 2.44 GiB | 40.9 (0.90) | 36.6 (0.84) | 32.2 (0.81) | 26.8 (0.62) | 20.5 (0.61) |
| Q8_0, quantized here | 3.85 GiB | 36.5 (0.84) | 32.6 (0.83) | 30.4 (0.78) | 29.0 (0.83) | 20.0 (0.76) |
| dzannotti BF16 | 7.24 GiB | 29.4 (0.77) | – | – | 25.5 (0.82) | – |
Decode in tokens per second, draft acceptance in brackets. This is where the 40 t/s from the short benchmarks come from — and where they go. The crossover sits between 8k and 30k tokens: the small head leads by 4.4 t/s at an empty context and trails by 2.2 t/s at 30k, because its acceptance holds up to 8k and then falls to 0.62 while the Q8_0 head stays near 0.83. The agent runs measure 21.9 to 24.4 t/s; their contexts are tens of thousands of tokens deep, so that is the expected figure.
At 64k the choice of head stops mattering: 20.5 against 20.0 t/s is noise. The Q8_0 head still drafts better there (0.76 against 0.61), but attention and indexer over 64k of context dominate every step, so the better hit rate buys no speed. Its advantage is a band around 30k, not a trend that keeps growing. Prompt processing drops from 318 to 261 tokens per second between 30k and 64k, so time to first token grows from 94 to 246 seconds. One caveat from the same runs: up to 30k every answer repeated the code word planted in the filler text, at 64k neither head did — one prompt shape, not a recall benchmark, but worth knowing before filling the window.
The Q8_0 head is not published. Quantize it from the BF16 head, then anything in state/mtp/ is picked
up automatically:
hf download dzannotti/Qwen3.8-Flash-Next-MTP-GGUF Qwen3.8-Flash-Next-MTP-BF16.gguf
engine/build-engramhalo/bin/llama quantize <BF16 file> state/mtp/Qwen3.8-Flash-Next-MTP-Q8_0.gguf Q8_0 8
Each point is a single measurement on one kind of task, so differences below about 1 t/s are noise.
Multi-user throughput (EngramHalo, 8 slots, 20k context per slot, 256 tokens per request)
| Users | Σ without MTP (IQ4_XS) | per user | TTFT | Σ with MTP (Q4_K_XL) | per user | Mix-ups |
|---|---|---|---|---|---|---|
| 1 | 20.4 | 21.4 | 0.6 s | 33.2 | 35.9 | 0 |
| 2 | 32.0 | 17.2 | 0.9 s | 30.8 | 16.9 | 0 |
| 4 | 42.4 | 11.6 | 1.8 s | 35.5 | 10.2 | 0 |
| 8 | 50.4 | 6.8 | 2.7 s | 34.8 | 5.1 | 0 |
With 8 users, continuous batching delivers 2.5 times the single throughput. MTP is only worth it with one user. The mix-up check (every user gets a code word) found no swapped slot with the #25992 patch.
Further measurements
- Hadamard rotation with q8_0 KV on vs. off (stock, Q2_K_XL): 24.2 vs. 24.6 tokens/s.
- Buffer sizes: compute 297 MB (ubatch 512) and 1188 MB (ubatch 2048); context 32k q8_0 = 673 MB; MTP draft 2146 MB + 740 MB.
- mmap load time on the stock fork with a warm page cache (Q2_K_XL): 14 s; with a cold cache (Q4_K_XL): 140 min.