Spark recipe

Models with a native DGX Spark recipe / support (Spark recipe tag). ← all tag kinds

Spark recipe (78)

ConfigurationEngineQuantCtxConcDecode tok/sStatusCompleted
Qwen3.8-27B · SGLang · NVFP4 + DSpark · conc-32 SGLang + DSpark NVFP4 262144 32 blocked 2026-08-20 05:25 +0800
Qwen3.8-27B · SGLang · NVFP4 + DSpark · conc-16 SGLang + DSpark NVFP4 262144 16 blocked 2026-08-20 05:25 +0800
Qwen3.8-27B · SGLang · NVFP4 + DFlash2 · conc-8 SGLang + DFlash2 NVFP4 262144 8 blocked 2026-08-20 05:25 +0800
Qwen3.8-27B · SGLang · NVFP4 + DFlash2 · conc-32 SGLang + DFlash2 NVFP4 262144 32 blocked 2026-08-20 05:25 +0800
Qwen3.8-27B · SGLang · NVFP4 + DFlash2 · conc-16 SGLang + DFlash2 NVFP4 262144 16 blocked 2026-08-20 05:25 +0800
Qwen3.8-27B · SGLang · NVFP4 + DFlash2 · conc-4 SGLang + DFlash2 NVFP4 262144 4 85.1 done 2026-08-20 05:24 +0800
Qwen3.8-27B · SGLang · NVFP4 + DFlash2 · conc-2 SGLang + DFlash2 NVFP4 262144 2 56.2 done 2026-08-20 05:16 +0800
Qwen3.8-27B · SGLang · NVFP4 + DFlash2 · conc-1 SGLang + DFlash2 NVFP4 262144 1 30.0 done 2026-08-20 05:04 +0800
Qwen3.8-27B · SGLang · NVFP4 + DSpark · conc-8 SGLang + DSpark NVFP4 262144 8 118.3 done 2026-08-20 04:46 +0800
Qwen3.8-27B · SGLang · NVFP4 + DSpark · conc-4 SGLang + DSpark NVFP4 262144 4 75.5 done 2026-08-20 04:17 +0800
Qwen3.8-27B · SGLang · NVFP4 + DSpark · conc-2 SGLang + DSpark NVFP4 262144 2 43.4 done 2026-08-20 03:46 +0800
Qwen3.8-27B · SGLang · NVFP4 + DSpark · conc-1 SGLang + DSpark NVFP4 262144 1 23.4 done 2026-08-20 03:31 +0800
Nemotron-3 Puzzle 75B-A9B · vLLM · NVFP4 · conc 32 vLLM NVFP4 65536 32 177.8 done 2026-07-11 17:21 +08
Nemotron-3 Puzzle 75B-A9B · vLLM · NVFP4 · conc 8 vLLM NVFP4 65536 8 94.0 done 2026-07-11 16:54 +08
Nemotron-3 Puzzle 75B-A9B · vLLM · NVFP4 · conc 1 vLLM NVFP4 65536 1 19.8 done 2026-07-11 16:29 +08
Qwen3.6-35B-A3B · vLLM-ultimate (AEON) · NVFP4 + DFlash · conc 128 (memory ceiling — watchdog trip) vLLM (aeon-vllm-ultimate custom container, v0.23.0+aeon.sm121a.dflash) + DFlash (z-lab/Qwen3.6-35B-A3B-DFlash @31977fbe small-page rev, num_speculative_tokens 11) NVFP4 (modelopt_mixed — W4A16_NVFP4 experts + FP8 GDN gates) 65536 128 425.7 done 2026-07-04 14:17 +0800
Qwen3.6-35B-A3B · vLLM-ultimate (AEON) · NVFP4 + DFlash · conc 64 vLLM (aeon-vllm-ultimate custom container, v0.23.0+aeon.sm121a.dflash) + DFlash (z-lab/Qwen3.6-35B-A3B-DFlash @31977fbe small-page rev, num_speculative_tokens 11) NVFP4 (modelopt_mixed — W4A16_NVFP4 experts + FP8 GDN gates) 65536 64 420.4 done 2026-07-04 14:11 +0800
Qwen3.6-35B-A3B · vLLM · NVFP4 · conc 128 vLLM NVFP4 65536 128 675.9 done 2026-07-04 13:50 +0800
Qwen3.6-35B-A3B · vLLM · NVFP4 · conc 64 vLLM NVFP4 65536 64 547.9 done 2026-07-04 13:41 +0800
Qwen3.6-35B-A3B · vLLM · NVFP4 + MTP · conc 128 (the knee) vLLM + MTP NVFP4 65536 128 749.9 done 2026-07-04 12:31 +0800
Qwen3.6-35B-A3B · vLLM · NVFP4 + MTP · conc 64 vLLM + MTP NVFP4 65536 64 670.8 done 2026-07-04 12:09 +0800
gpt-oss-20b · vLLM · MXFP4 · conc 8 vLLM MXFP4 65536 8 212.5 done 2026-07-02 07:51 +0800
gpt-oss-20b · vLLM · MXFP4 · conc 1 vLLM MXFP4 65536 1 45.6 done 2026-07-02 07:51 +0800
Qwen3.6-35B-A3B · vLLM-ultimate (AEON) · NVFP4 + DFlash · conc 1 vLLM (aeon-vllm-ultimate custom container, v0.23.0+aeon.sm121a.dflash) + DFlash (z-lab/Qwen3.6-35B-A3B-DFlash @31977fbe small-page rev, num_speculative_tokens 11) NVFP4 (modelopt_mixed — W4A16_NVFP4 experts + FP8 GDN gates) 65536 1 99.8 done 2026-07-02 07:21 +0800
Qwen3.6-35B-A3B · vLLM-ultimate (AEON) · NVFP4 + DFlash · conc 8 vLLM (aeon-vllm-ultimate custom container, v0.23.0+aeon.sm121a.dflash) + DFlash (z-lab/Qwen3.6-35B-A3B-DFlash @31977fbe small-page rev, num_speculative_tokens 11) NVFP4 (modelopt_mixed — W4A16_NVFP4 experts + FP8 GDN gates) 65536 8 267.4 done 2026-07-02 07:21 +0800
Qwen3.6-35B-A3B · vLLM-ultimate (AEON) · NVFP4 + DFlash · conc 4 vLLM (aeon-vllm-ultimate custom container, v0.23.0+aeon.sm121a.dflash) + DFlash (z-lab/Qwen3.6-35B-A3B-DFlash @31977fbe small-page rev, num_speculative_tokens 11) NVFP4 (modelopt_mixed — W4A16_NVFP4 experts + FP8 GDN gates) 65536 4 210.7 done 2026-07-02 07:21 +0800
Qwen3.6-35B-A3B · vLLM-ultimate (AEON) · NVFP4 + DFlash · conc 32 vLLM (aeon-vllm-ultimate custom container, v0.23.0+aeon.sm121a.dflash) + DFlash (z-lab/Qwen3.6-35B-A3B-DFlash @31977fbe small-page rev, num_speculative_tokens 11) NVFP4 (modelopt_mixed — W4A16_NVFP4 experts + FP8 GDN gates) 65536 32 407.1 done 2026-07-02 07:21 +0800
Qwen3.6-35B-A3B · vLLM-ultimate (AEON) · NVFP4 + DFlash · conc 2 vLLM (aeon-vllm-ultimate custom container, v0.23.0+aeon.sm121a.dflash) + DFlash (z-lab/Qwen3.6-35B-A3B-DFlash @31977fbe small-page rev, num_speculative_tokens 11) NVFP4 (modelopt_mixed — W4A16_NVFP4 experts + FP8 GDN gates) 65536 2 151.3 done 2026-07-02 07:21 +0800
Qwen3.6-35B-A3B · vLLM-ultimate (AEON) · NVFP4 + DFlash · conc 16 vLLM (aeon-vllm-ultimate custom container, v0.23.0+aeon.sm121a.dflash) + DFlash (z-lab/Qwen3.6-35B-A3B-DFlash @31977fbe small-page rev, num_speculative_tokens 11) NVFP4 (modelopt_mixed — W4A16_NVFP4 experts + FP8 GDN gates) 65536 16 340.2 done 2026-07-02 07:21 +0800
Qwen3-Coder-30B-A3B · DDTree vs DFlash · HumanEval (coding workload) · single-stream DDTree research harness (github.com/liranringel/ddtree, PyTorch + transformers, batch-1) + DDTree tree draft (budgets 64/256) vs single-line DFlash — z-lab/Qwen3-Coder-30B-A3B-DFlash, block_size 16 BF16 (harness loads the unquantized target via AutoModelForCausalLM) 4096 1 49.3 done 2026-07-01 21:26 +0800
Qwen3-Coder-30B-A3B · DDTree vs DFlash vs autoregressive · single-stream (batch-1) DDTree research harness (github.com/liranringel/ddtree, PyTorch + transformers, batch-1) + DDTree tree draft (budgets 64/256) vs single-line DFlash — z-lab/Qwen3-Coder-30B-A3B-DFlash, block_size 16 BF16 (harness loads the unquantized target via AutoModelForCausalLM) 4096 1 20.8 done 2026-07-01 21:13 +0800
gpt-oss-120b · vLLM · MXFP4 + EAGLE3 (LMSYS draft) · conc 8 vLLM + EAGLE3 (lmsys/EAGLE3-gpt-oss-120b-bf16 — the SGLang/SpecForge draft, on vLLM) MXFP4 65536 8 175.3 done 2026-07-01 19:40 +0800
gpt-oss-120b · vLLM · MXFP4 + EAGLE3 (LMSYS draft) · conc 32 vLLM + EAGLE3 (lmsys/EAGLE3-gpt-oss-120b-bf16 — the SGLang/SpecForge draft, on vLLM) MXFP4 65536 32 246.7 done 2026-07-01 19:21 +0800
gpt-oss-120b · vLLM · MXFP4 + EAGLE3 (LMSYS draft) · conc 1 vLLM + EAGLE3 (lmsys/EAGLE3-gpt-oss-120b-bf16 — the SGLang/SpecForge draft, on vLLM) MXFP4 65536 1 25.5 done 2026-07-01 17:37 +0800
gpt-oss-20b · vLLM · MXFP4 + EAGLE3 · conc 16 vLLM + EAGLE3 MXFP4 65536 16 431.8 done 2026-07-01 16:36 +0800
gpt-oss-20b · vLLM · MXFP4 + EAGLE3 · conc 4 vLLM + EAGLE3 MXFP4 65536 4 85.4 done 2026-07-01 16:15 +0800
gpt-oss-20b · vLLM · MXFP4 + EAGLE3 · conc 2 vLLM + EAGLE3 MXFP4 65536 2 59.0 done 2026-07-01 16:02 +0800
gpt-oss-20b · vLLM · MXFP4 · conc 16 vLLM MXFP4 65536 16 340.7 done 2026-07-01 15:35 +0800
gpt-oss-20b · vLLM · MXFP4 · conc 4 vLLM MXFP4 65536 4 127.3 done 2026-07-01 15:22 +0800
gpt-oss-20b · vLLM · MXFP4 · conc 2 vLLM MXFP4 65536 2 83.4 done 2026-07-01 15:22 +0800
Qwen3.6-35B-A3B · vLLM · NVFP4 · conc 16 vLLM NVFP4 65536 16 332.4 done 2026-07-01 14:58 +0800
Qwen3.6-35B-A3B · vLLM · NVFP4 · conc 8 vLLM NVFP4 65536 8 241.8 done 2026-07-01 14:39 +0800
Qwen3.6-35B-A3B · vLLM · NVFP4 · conc 4 vLLM NVFP4 65536 4 173.8 done 2026-07-01 14:24 +0800
Qwen3.6-35B-A3B · vLLM · NVFP4 · conc 2 vLLM NVFP4 65536 2 113.2 done 2026-07-01 14:07 +0800
Qwen3.6-35B-A3B · vLLM · NVFP4 · conc 1 vLLM NVFP4 65536 1 74.7 done 2026-07-01 13:50 +0800
Qwen3.6-35B-A3B · vLLM · NVFP4 + MTP · conc 16 vLLM + MTP NVFP4 65536 16 433.3 done 2026-07-01 12:54 +0800
Qwen3.6-35B-A3B · vLLM · NVFP4 + MTP · conc 4 vLLM + MTP NVFP4 65536 4 232.4 done 2026-07-01 12:37 +0800
Qwen3.6-35B-A3B · vLLM · NVFP4 + MTP · conc 2 vLLM + MTP NVFP4 65536 2 161.2 done 2026-07-01 12:20 +0800
Qwen3.6-35B-A3B · DDTree (Diffusion Draft Tree) · BLOCKED (hybrid GDN cache) DDTree research harness (github.com/liranringel/ddtree, PyTorch + transformers, batch-1) + DDTree — tree draft over the z-lab/Qwen3.6-35B-A3B-DFlash block-diffusion drafter BF16 (harness loads the unquantized target via AutoModelForCausalLM) 4096 1 blocked 2026-07-01
Qwen3.6-27B (dense) · DDTree (Diffusion Draft Tree) · BLOCKED (hybrid GDN cache) DDTree research harness (github.com/liranringel/ddtree, PyTorch + transformers, batch-1) + DDTree — tree draft over the z-lab/Qwen3.6-27B-DFlash block-diffusion drafter BF16 (harness loads the unquantized target via AutoModelForCausalLM) 4096 1 blocked 2026-07-01
DeepSeek V4-Flash REAP-K180 · ds4 · NVFP4-hybrid · conc 1 (1M ctx) ds4 (ds4-nvfp4-spark, custom GB10 runtime) + none — ds4 is a serial single-worker; the DeepSeek MTP head is not exercised by this runtime NVFP4-hybrid (NVFP4 experts + Q2_K + Q8_0 + F16) 1048576 1 11.4 done 2026-06-28 13:34 +08
MiniMax-M2.7-REAP 172B · vLLM · NVFP4 (W4A4) vLLM NVFP4 (W4A4) 65536 32 111.9 done 2026-06-27 19:22 +0800
MiniMax-M2.7-REAP 172B · vLLM · NVFP4 (W4A4) · conc 1 · 160K vLLM NVFP4 (W4A4) 163840 1 25.4 done 2026-06-27 19:22 +0800
MiniMax-M2.5-REAP 139B · vLLM · NVFP4 (W4A16) · conc 1 · 192K vLLM NVFP4 (W4A16) 196608 1 27.0 done 2026-06-27 15:22 +0800
MiniMax-M2.5-REAP 139B · vLLM · NVFP4 (W4A16) vLLM NVFP4 (W4A16) 65536 32 120.0 done 2026-06-26 22:25 +0800
GLM-4.5-Air-REAP 82B · vLLM · NVFP4 (W4A16) vLLM NVFP4 (W4A16) 65536 32 158.4 done 2026-06-26 16:10 +0800
gpt-oss-120b · SGLang · MXFP4 + EAGLE3 · conc 32 SGLang + EAGLE3 MXFP4 65536 32 171.9 done 2026-06-23 21:20 +0800
gpt-oss-120b · SGLang · MXFP4 + EAGLE3 · conc 1 SGLang + EAGLE3 MXFP4 65536 1 40.6 done 2026-06-23 20:08 +0800
Qwen3.6-35B-A3B · vLLM · NVFP4 + MTP · conc 1 vLLM + MTP NVFP4 65536 1 93.9 done 2026-06-23 13:02 +08
Qwen3.6-35B-A3B · vLLM · NVFP4 + MTP · conc 8 vLLM + MTP NVFP4 65536 8 289.1 done 2026-06-23 12:49 +08
Qwen3.6-35B-A3B · vLLM · NVFP4 + MTP vLLM + MTP NVFP4 65536 32 541.3 done 2026-06-23 12:34 +08
Qwen3.6-27B · SGLang · NVFP4 + MTP SGLang + MTP (NEXTN) NVFP4 65536 32 196.6 done 2026-06-23 12:18 +08
Qwen3.6-35B-A3B · vLLM · NVFP4 vLLM NVFP4 65536 32 430.8 done 2026-06-23 11:13 +08
gpt-oss-120b · vLLM · MXFP4 + EAGLE3 · conc 8 vLLM + EAGLE3 MXFP4 65536 8 50.0 done 2026-06-22 23:44 +08
gpt-oss-120b · vLLM · MXFP4 + EAGLE3 · conc 1 vLLM + EAGLE3 MXFP4 65536 1 14.7 done 2026-06-22 23:29 +08
gpt-oss-20b · vLLM · MXFP4 + EAGLE3 · conc 8 vLLM + EAGLE3 MXFP4 65536 8 126.3 done 2026-06-22 23:13 +08
gpt-oss-20b · vLLM · MXFP4 + EAGLE3 · conc 1 vLLM + EAGLE3 MXFP4 65536 1 38.6 done 2026-06-22 23:03 +08
gpt-oss-20b · vLLM · MXFP4 + EAGLE3 vLLM + EAGLE3 MXFP4 65536 32 686.5 done 2026-06-22 17:52 +08
gpt-oss-120b · vLLM · MXFP4 + EAGLE3 vLLM + EAGLE3 MXFP4 65536 32 138.5 done 2026-06-22 17:42 +08
gpt-oss-120b · vLLM · MXFP4 vLLM MXFP4 65536 32 252.8 done 2026-06-22 15:30 +08
gpt-oss-20b · vLLM · MXFP4 vLLM MXFP4 65536 32 535.3 done 2026-06-22 06:00 +08
Nemotron-3 Nano-4B · llama.cpp · Q4_K_M llama.cpp Q4_K_M 65536 32 141.3 done 2026-06-22 03:10 +08
Nemotron-3 Super 120B · vLLM · NVFP4 vLLM NVFP4 65536 32 96.9 done 2026-06-22 00:22 +08
Nemotron-3 Nano-Omni 30B-A3B · vLLM · NVFP4 vLLM NVFP4 65536 32 388.9 done 2026-06-21 23:43 +08
Nemotron-3 Elastic 30B-A3B · vLLM · NVFP4 vLLM NVFP4 65536 32 353.0 done 2026-06-21 23:23 +08
gpt-oss-120b · SGLang · MXFP4 SGLang MXFP4 65536 32 140.3 done 2026-06-21 22:56 +08
gpt-oss-20b · SGLang · MXFP4 SGLang MXFP4 65536 32 279.5 done 2026-06-21 22:08 +08
Qwen3.6-27B (dense) · DDTree (Diffusion Draft Tree) · BLOCKED (hybrid GDN cache) DDTree research harness (github.com/liranringel/ddtree, PyTorch + transformers, batch-1) + DDTree — tree draft over the z-lab/Qwen3.6-27B-DFlash block-diffusion drafter BF16 (harness loads the unquantized target via AutoModelForCausalLM) 4096 1 blocked 2026-07-01
Qwen3.6-35B-A3B · DDTree (Diffusion Draft Tree) · BLOCKED (hybrid GDN cache) DDTree research harness (github.com/liranringel/ddtree, PyTorch + transformers, batch-1) + DDTree — tree draft over the z-lab/Qwen3.6-35B-A3B-DFlash block-diffusion drafter BF16 (harness loads the unquantized target via AutoModelForCausalLM) 4096 1 blocked 2026-07-01
Qwen3.8-27B · SGLang · NVFP4 + DFlash2 · conc-16 SGLang + DFlash2 NVFP4 262144 16 blocked 2026-08-20 05:25 +0800
Qwen3.8-27B · SGLang · NVFP4 + DFlash2 · conc-32 SGLang + DFlash2 NVFP4 262144 32 blocked 2026-08-20 05:25 +0800
Qwen3.8-27B · SGLang · NVFP4 + DFlash2 · conc-8 SGLang + DFlash2 NVFP4 262144 8 blocked 2026-08-20 05:25 +0800
Qwen3.8-27B · SGLang · NVFP4 + DSpark · conc-16 SGLang + DSpark NVFP4 262144 16 blocked 2026-08-20 05:25 +0800
Qwen3.8-27B · SGLang · NVFP4 + DSpark · conc-32 SGLang + DSpark NVFP4 262144 32 blocked 2026-08-20 05:25 +0800
gpt-oss-20b · llama.cpp · MXFP4 llama.cpp MXFP4 131072 32 blocked 2026-06-21