The short answer?
DeepSeek V4 Flash runs on 2× DGX Spark with TP=2, DSpark speculative decoding, NVFP4 KV cache, and 1M context. The reference recipe is MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark with image ghcr.io/anemll/dspark-vllm-gx10:0.1.1. Four hotfixes are mandatory: disable earlyoom, spin-wait P-core (#79), assistant-final hotfix (#52), and tool-call truncation (#55). Expected speed: 62–83 tok/s for one chat ≤128K, and ~160–190 tok/s aggregate for six short chats — expected numbers from the README, to be re-measured on our own hardware.
📊 Sources: README + CHANGELOG of the MiaAI-Lab recipe (777★, commit 18/08/2026), and 5ac's own research on 2× DGX Spark DeepSeek V4 Flash. Throughput figures are "expected" from Anemll, not yet measured on 5ac hardware.
Why DeepSeek V4 Flash needs 2× DGX Spark
DeepSeek V4 Flash is a Mixture of Experts model of roughly 284 billion parameters. A single DGX Spark node has 128GB of unified memory (CPU and GPU shared, CUDA sees roughly 121.7 GiB). Running a model of this size on one node — especially with true 1M context and NVFP4 KV cache — is not feasible: the combined weights and KV cache exceed memory, and decode is capped by memory bandwidth (~273 GB/s) rather than by compute.
Two nodes unlock three things at once. First, tensor parallelism (TP=2): the model is split across the two machines, each holding part of the weights and participating in every inference step. Second, DSpark speculative decoding: predict several tokens and verify once, which wins big on bandwidth-bound hardware like GB10. Third, true 1M context with 0.835 text utilization — about 2.49 million tokens, enough to hold a large business's working context.
But two nodes means twice the complexity: synchronization, network fabric, NCCL, and the order in which workers and head boot. This is where a mature recipe is worth far more than building your own.
Picking the recipe: MiaAI-Lab, not "build it yourself to look cool"
After scanning the open-source ecosystem, we settled on MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark (777★, latest commit 18/08/2026, CI validates every push) as our base. Not because it is popular, but because it already ships a set of hotfixes that a self-built version would take a long time to discover on its own.
One honest caveat: we use only the serving config from this repo, not the Hugging Face models from MiaAI-Lab (Qwable/Gemmable). A public discussion raised a claim that "Qwable" is actually a quantized/dequantized base model rather than a real fine-tune — we are not taking that risk. The serving config and code, on the other hand, are measured in practice, credit sources clearly, and are trustworthy.
The base image: ghcr.io/anemll/dspark-vllm-gx10:0.1.1
The base image is ghcr.io/anemll/dspark-vllm-gx10:0.1.1 (from Anemll/dspark-vllm-gx10), built specifically for GB10 (ARM64/aarch64). This is a key point: GB10 cannot run a standard x86 image, and this image has been forked, audited, and shipped with the latest hotfixes.
For 5ac, this image is better than the aidendle94/sparkrun-vllm-ds4-gb10 image previously noted in our 2× Spark purchase note — we are updating our deploy decision accordingly. The checkpoint is the official 0731 (ABLITERATED=0), meaning we do not use abliterated weights.
Key takeaway: With 2-node hardware, do not try to build the image yourself. The official Anemll image already packages the right libraries for SM121 (no native FP4 tensor cores, so NVFP4 runs through the flashinfer/triton backend). Building x86 or using a wrong-architecture image will cost a day of NCCL debugging.
Default profile configuration
The recipe's default configuration (full table below) is the safest starting point. Do not tune before you measure:
| Setting | Value |
|---|---|
| Image | ghcr.io/anemll/dspark-vllm-gx10:0.1.1 |
| Checkpoint | official 0731 (ABLITERATED=0) |
| Context | MAX_MODEL_LEN=1048576 (1M) |
| Concurrency | MAX_NUM_SEQS=6 · MAX_NUM_BATCHED_TOKENS=8192 |
| KV cache | nvfp4_ds_mla, text util 0.835 (~2.49M tokens) |
| Speculative | MTP_NUM_TOKENS=5 (≥ dspark_block_size) |
| Thinking | DEFAULT_THINKING=max · VLLM_USE_BREAKABLE_CUDAGRAPH=0 |
Two values stand out. --max-cudagraph-capture-size should equal MAX_NUM_SEQS × (MTP_NUM_TOKENS+1) = 36 (6×6). And the TP=2 network env is WORKER_HOST, MASTER_ADDR, NCCL_IB_HCA, NCCL_SOCKET_IFNAME, VLLM_HOST_IP, HF_CACHE. The recipe's resolve_nccl_gid_indexes() resolves the GID automatically at every launch — preventing NCCL failures after reboot. The link between the two nodes must be a QSFP56 RoCE 200G DAC cable (e.g. Mellanox MCP1650-H00AE30), not ordinary Ethernet.
Four mandatory hotfixes — the difference between "demo" and "production"
This is the most valuable part of the recipe. The four hotfixes below are what turn a "demo works" cluster into a "production works" one. If you run agentic workloads, missing any of them costs you downtime or a poisoned transcript.
1. Disable earlyoom on both hosts. earlyoom kills processes based on memory — under deep-context load it will kill vLLM the moment memory climbs, even though vLLM is working correctly. It is the most confusing source of interruption and the first thing to turn off.
2. Spin-wait P-core (#79). Lower SpinCondition.busy_loop_s from 1 to 0.002 on TP=2. The old version made spin-wait hog an entire Grace P-core, wasting your most expensive resource. There is an opt-out DSPARK_SKIP_SPIN_WAIT_HOTFIX=1 if you need the old behavior.
3. Assistant-final hotfix (#52/PR#53). Fixes the agent-harness loop: when a request ends with an assistant message, the stock build returns a bare EOS with no generation header → the agent loop produces no-op turns and hallucinated DSML. It is now opt-in via DSPARK_ENABLE_ASSISTANT_FINAL_HOTFIX=1 (default 0 = stock), fail-closed (restores bytes on failure) and idempotent.
4. Tool-call truncation (#55). Previously, when max_tokens cut mid tool call, the model reported finish_reason="length" with "tool_calls" plus truncated JSON — poisoning the transcript with a 400. This hotfix prevents exactly that. Note the protocol limit: streaming args remain append-only.
Key takeaway: Hotfixes #21/#26/#27/#43 always run and cannot be skipped by DSPARK_SKIP_HOTFIX; #22 is separate. And remember: VLLM_DSPARK_* env vars are no-ops on Anemll 0.1.1 (they only take effect if you switch to a Stage-C image).
Beyond the four mandatory ones, the recipe ships other useful hotfixes: start after reboot (#72) — when the container is already up it returns exit 3 instead of exit 1, so systemd uses SuccessExitStatus=3 to avoid a retry storm; RULER-lite context (#81) — bulk-pads the right 32k/262k cells (previously capped at ~4.8k even on exit 0); thinking budget (#31/#48) — two Triton kernels hard-cap reasoning, raising speed from 37.2 to 59.3 tok/s and removing the ~4.5 tok/s cliff.
Ops scripts and systemd after reboot
The recipe bundles a full set of ops scripts: start- / stop- / status- / logs- / smoke-*.sh, run in worker-first, head-last order. This matters for TP=2: the worker node must be ready before head launches, or the NCCL rendezvous hangs.
For production, wire the scripts into systemd so the cluster restarts safely after reboot. Because hotfix #72 already returns exit 3 when the container is up, set SuccessExitStatus=3 so systemd does not retry endlessly. Combined with resolve_nccl_gid_indexes() auto-resolving the GID, your cluster survives reboots without manual intervention.
For extra capability, the recipe has an optional ENABLE_VL_SIDECAR=1 (Qwen3-VL on :8889 + MCP) and ENABLE_VLLM_GB10_PATCH=1 (--quantization modelopt_gb10_hybrid, default OFF).
Responses API live verifier — 4 gates
What sets this recipe apart from a self-build: the verify-responses-api-live.py script directly checks four gates of the Responses API:
- Text/SSE — streaming responses in the correct format.
- Stateful tool continuation — the agent calls a tool and continues with correct context (requires
VLLM_ENABLE_RESPONSES_API_STORE=1). - Strict JSON schema — output matches the declared schema.
- Reasoning + multi-turn prefix reuse + disconnect cleanup — long-running behavior and cleanup when a client disconnects.
Running this verifier after every deploy is the fastest way to know whether the cluster is "alive in the real sense" or merely "reachable on the network". For agentic workloads, these four gates are what you need to trust before putting agents into production.
Tuning applied from GB10 benchmarks
Once the cluster is stable, these tunings were measured on real GB10 (from the Qwen3.8-27B-SGLang repo and 5ac's internal benchmarks) and apply to the way we operate DeepSeek V4 Flash:
- GDN state bf16 instead of float32 — float32 is about 3% slower.
- Spec steps 3/1/4 is the peak. A sweep 2→6 gives 12.8 / 17.2 / 16.8 / 16.3 / 15.8 tok/s thinking; 3/1/4 is highest.
- Pin 10 Cortex-X5 cores (
--cpuset-cpus 5-9,15-19) → +2–7% decode, since the scheduler/tokenizer no longer floats to the A725 at 2.8GHz. - Mem 0.90 · chunked prefill 8192 · FP8 KV cache (
fp8_e4m3) for models using FP8.
One limitation to remember: DSpark does not use YaRN / context > 262K (a build limitation). If you truly need 1M context, that is exactly what the NVFP4 + MTP configuration of this recipe is doing — which is why DeepSeek V4 Flash on 2× Spark is not bound by that limit.
FAQ
Can this recipe run on a single DGX Spark?
Not for DeepSeek V4 Flash. The model requires TP=2 with 1M context and NVFP4 — 2× Spark is the minimum viable configuration. If you only have one node, look at a lighter model (like Qwen3.8-27B FP8) that fits in 128GB.
Are the 62–83 tok/s figures measured on 5ac hardware?
Not yet. They are expected numbers from the Anemll README, not measurements on 5ac's hardware. We will re-measure n=5 when the machines arrive and publish the results on the blog. In this article, every number marked "expected" follows that convention.
Do I need a special RoCE cable between the 2 nodes?
Yes. A QSFP56 RoCE 200G DAC cable is mandatory for TP=2 to run at full speed (e.g. NVIDIA/Mellanox MCP1650-H00AE30). Running over ordinary Ethernet will slow NCCL and the cluster will not reach expected throughput.
Do agentic workloads need the assistant-final hotfix?
Yes, effectively mandatory. If a request ends with an assistant message, the stock build (default 0) will produce no-op turns and hallucinated DSML. Enable DSPARK_ENABLE_ASSISTANT_FINAL_HOTFIX=1 so the agent-harness loop runs correctly.
Conclusion
Deploying DeepSeek V4 Flash on 2× DGX Spark is not about buying machines and running one command. It is about choosing the right recipe, using the right image, applying every hotfix, and trusting a verifier that measures reality instead of "it answered, so it works".
For 5ac, the MiaAI-Lab/Anemll recipe is the most mature starting point we found in the open-source ecosystem: full ops scripts, fail-closed hotfixes, and a Responses API verifier. We forked it, audited it, and will re-measure it on our own hardware — because every recipe in the 5ac repo runs for real, not on a slide.