The takeaway?
DSpark and MTP are both speculative decoding, but neither wins outright on the DGX Spark (GB10). Measured n=5 on the same NVFP4 weights: DSpark wins code (51.5 vs 34.5 tok/s) and short chat (23.2 vs 21.0); MTP wins long essays (24.1 vs 18.3 tok/s). For agentic, coding and tool-calling workloads, DSpark is the right choice. One limit to remember: DSpark does not use YaRN, so context stops at roughly 262K tokens.
📊 Sources: DSpark vs MTP A/B by MiaAI-Lab measured on a real GB10 (18/08/2026, n=5, same NVFP4 weights); FP8/NVFP4 + MTP benchmark by Kubesimplify (14–17/08/2026). Sources detailed at the end.
If you have ever wondered why a powerful AI box like the DGX Spark still breathes hard while generating one token at a time, the answer is memory bandwidth. The GB10 has about 273 GB/s of bandwidth, but every token it produces must read the entire set of model weights. Decoding is therefore bound by bandwidth, not by compute. That is exactly why speculative decoding makes a big difference on this kind of hardware.
What speculative decoding is, and why it matters on the GB10
The idea is simple: instead of letting the large model generate one token at a time, you use a lighter prediction mechanism to guess several tokens at once, then have the large model verify the whole block in a single pass. When the prediction is mostly right, you save many full-weight reads. On a chip bound by bandwidth like the GB10, this is the most effective way to raise token/s without changing hardware.
In the Qwen3.8-27B ecosystem running on the GB10 with SGLang, two mechanisms are common: DSpark and MTP. They share the same goal but differ in how they predict and how they run, which leads to different results by workload. We do not trust marketing, so we ran an A/B to measure.
The A/B setup: n=5, same weights, only the mechanism differs
To make the result credible, we kept everything else identical and changed only one variable: the speculative decoding mechanism. Same NVFP4 weights, same Qwen3.8-27B model, same serving configuration, and each workload run 5 times (n=5).
- DSpark — run through the
start-dspark.shscript, block-7 configuration. - MTP — run through the
start.shscript, EAGLE 3/1/4 configuration (Multi-Token Prediction).
Three workload types were measured, standing in for three real problem groups: generated code with tests, short chat with thinking, and long free-form text.
| Workload | DSpark (block-7) | MTP (EAGLE 3/1/4) | Winner |
|---|---|---|---|
| Code — LRUCache + test | 51.5 tok/s (51.4–51.7) | 34.5 tok/s (34.5–34.6) | DSpark ~1.5× |
| Short chat (thinking on) | 23.2 tok/s | 21.0 tok/s | DSpark slightly |
| Long essay (Babbage → GPUs) | 18.3 tok/s | 24.1 tok/s | MTP ~1.3× |
Source: MiaAI-Lab, Qwen3.8-27B-SGLang-DGX-Spark repo, measured 18/08/2026 on a GB10, n=5, same NVFP4 weights.
Analysis: agents and code win with DSpark, long essays with MTP
DSpark wins two of the three workloads, and wins the code problem most clearly — 51.5 versus 34.5 tok/s. This matters to 5ac, because our main workload is agentic: code, tool-calling, and repeating chains of actions. DSpark tracks the token distribution closely when code has tight structure, so its drafter predicts correctly more often.
Short chat with thinking is a close match: 23.2 versus 21.0 tok/s. DSpark still wins, but the edge is thin. In reasoning mode the token stream is harder to predict, so the advantage of speculative decoding narrows.
The key point: long essays are the exception. When generating long free-form text, MTP reaches 24.1 tok/s versus 18.3 for DSpark — about a 1.3× edge. If your workload is long-form writing, MTP deserves consideration.
The source conclusion is concise: "Use DSpark for agents, code and normal chat; MTP only if you write long essays." For 5ac — whose workload is mainly agentic, coding and tool-calling — DSpark is the right choice. This also matches our deployment decision for Qwen3.8 on the GB10: FP8 + DFlash, stable, no crashes.
Cross-checking against Kubesimplify: FP8, NVFP4 and MTP
These results do not stand alone. Kubesimplify's benchmark (measured on a GB10, 14–17/08/2026) with Qwen3.8-27B adds a wider picture of the whole quantization and decode ecosystem, beyond the single DSpark vs MTP pair:
- BF16 (native): ~3–4 tok/s — slowest, for reference only.
- FP8 (vLLM, no spec): 8.2 tok/s single, 57.9 tok/s aggregate at 10 streams — clean and stable.
- FP8 + DFlash: 17.2 tok/s free-gen, 45+ tok/s edit, 65.6 tok/s aggregate — our stable pick.
- NVFP4 (vLLM): 11.5 tok/s single, 84.3 tok/s aggregate — fast but needs the Marlin backend.
- NVFP4 + MTP: 22.0 tok/s single, 105.8 tok/s aggregate — the fastest peak, yet it hard-rebooted the machine twice under long context and concurrency.
- Ollama Q4 + MTP: 26.5 tok/s single — fastest for a single stream.
Reading this table together with the DSpark vs MTP A/B, one lesson cuts across both: peak numbers and stability are two different things. NVFP4 + MTP delivers the highest aggregate speed (105.8 tok/s), yet it is the very configuration that caused a machine reboot. When your workload is an agent running continuously, one mid-shift crash costs more than the few tok/s difference on a table.
Measured tuning: spec steps, X5 core pinning and GDN bf16
Choosing the right mechanism is only half the job. The other half is tuning to push it to its peak. Three tuning steps measured on the GB10 are worth applying:
Spec steps 3/1/4 is the peak. The MTP head is trained for 3 steps, and the 3/1/4 configuration gives the best result. Sweeping from 2 to 6 steps gave 12.8 / 17.2 / 16.8 / 16.3 / 15.8 tok/s for thinking — 17.2 tok/s (the 3/1/4 config) is the peak.
Pin 10 Cortex-X5 cores. Use --cpuset-cpus 5-9,15-19 to pin the scheduler and tokenizer to 10 X5 cores, so they do not contend with other performance cores. Result: decode throughput rises 2–7%.
GDN state bf16. The cookbook defaults to float32, but moving the GDN state to bf16 is about 3% faster. Alongside this, decode measured on the GB10 reaches 17.2–20.5 tok/s (thinking) and 21.6–22.7 tok/s (non-thinking), with a TTFT of about 8.3 seconds for a warmed 16K prompt.
When to avoid DSpark: the 262K context limit
One important note you cannot skip. DSpark in this build does not use YaRN, and context is capped near ~262K tokens. If you need a 1-million-token context — for example processing long documents or huge customer records — you must switch to MTP.
A real trap: enabling YARN=1 with DSpark throws AttributeError: 'PreTrainedConfig' object has no attribute 'max_position_embeddings'. It does not quietly work at long context; it fails loudly, and you want to know that before you spend time debugging.
There is also a stability lesson worth recording. In Kubesimplify's benchmark (14–17/08/2026), the NVFP4 + MTP configuration caused the machine to hard-reboot twice under long-context and high concurrency. This is why we are cautious with MTP and prefer DSpark/DFlash for high-reliability workloads, even though MTP's peak speed can be higher.
Our recommendation for 5ac: agentic and coding → DSpark/DFlash
Based on this benchmark, our deployment decision for Qwen3.8 on the GB10 is clear: FP8 + DFlash speculative decoding, running on stable vLLM, no MTP. The reason is not only speed, but stability and a match with the real workload.
- Right workload: agentic, coding, tool-calling — where DSpark wins.
- Stable: DFlash runs on stable vLLM without crashes; MTP on vLLM nightly has hard-rebooted the machine twice (source: Kubesimplify).
- Simple: no SM121 FP4 gotcha — FP8 runs natively clean, with blockwise quality loss under 1%.
Measured speed of FP8 + DFlash: 17.2 tok/s for free generation and 45+ tok/s for edit cases, with ~65.6 tok/s aggregate at 10 concurrent streams. Compared to plain FP8 (~8.2 tok/s), DFlash delivers roughly 2–5× the speed.
Conclusion: there is no one-size-fits-all, measure your workload
The lesson from this benchmark is not "DSpark is better than MTP." It is: speculative decoding must be chosen by workload, and measured on your own hardware. On the same GB10, code wins with DSpark while long essays win with MTP. The numbers in a table only matter if you know which problem group you are running.
For 5ac the answer is clear: our main workload is agentic and coding, so we settled on DSpark/DFlash and reserve MTP for long-form writing if it comes up. Every number above was measured on a real GB10, and you can rerun the whole thing with our open recipe.
If you run a local AI cluster and are torn between the two mechanisms, our advice is direct: fork the recipe, run an n=5 benchmark with your own workload, and only then trust the numbers. Do not choose from a printed spec sheet.
What about you? Is your main workload code, chat, or long-form text? Leave a comment or fork our repo — we read everything and will share more hands-on GB10 tuning experience in the posts ahead.