Qwen3.6-35B-A3B · vLLM-ultimate (AEON) · NVFP4 + DFlash · conc 16
status: done
Configuration
| Model | nvidia/Qwen3.6-35B-A3B-NVFP4 |
|---|---|
| Company | Alibaba |
| Family | Qwen |
| Parameters | 35B / 3B (MoE, hybrid GDN+full-attn) + DFlash external drafter |
| Engine | vLLM (aeon-vllm-ultimate custom container, v0.23.0+aeon.sm121a.dflash) + DFlash (z-lab/Qwen3.6-35B-A3B-DFlash @31977fbe small-page rev, num_speculative_tokens 11) (speculative decoding) |
| Quant / precision | NVFP4 (modelopt_mixed — W4A16_NVFP4 experts + FP8 GDN gates) |
| Why this quant | conc-16 fine-grained point of the DFlash money-chart line, tracing the 8→32 batch-fill collapse where DFlash's wasted draft compute sinks aggregate throughput. REVISED 2026-07-02 — re-run at ctx 65536 (was 40960) as part of the full ctx-matched six-point re-sweep. Same one-boot protocol (official checkpoint, small-page drafter, max-num-seqs 64). SAFETY — untrusted third-party image; NO credentials, weights + drafter READ-ONLY, port loopback-only. |
| Download | nvidia/Qwen3.6-35B-A3B-NVFP4 |
| Context window | 65536 |
| Input modalities | text, image, video (served text-only here) |
Measured results
| Prefill tok/s | 355.12 |
|---|---|
| Decode tok/s | 340.2 |
| Peak memory (GB) | 111.3 (system MemAvailable delta (idle baseline 118 GiB → 6.7 GiB available at load) over the one-boot conc-1/2/4/8/16/32 sweep — vLLM static KV (util 0.85) + DFlash drafter) |
| Completed | 2026-07-02 07:21 +0800 |
Full run command
# UNTRUSTED image — NO creds; weights + drafter READ-ONLY; loopback port. ONE boot (ctx 65536,
# max-num-seqs 64, small-page drafter @31977fbe) sweeping client conc 1/2/4/8/16/32 — see the main
# (c1) page for the full docker run command. Only the client --concurrency differs.
python3 scripts/bench-serving.py --base-url http://127.0.0.1:8000 --model official \
--dataset benchmark_data/ShareGPT_V3_unfiltered_cleaned_split.json --concurrency 16 --num-prompts 1000 --max-seconds 600 --max-tokens 256
# 810/1000 prompts (hit 600 s cap), 0 errors.
Notes
conc-16 point of the Qwen3.6-35B-A3B NVFP4 + DFlash line — the batch-fill collapse, REVISED 2026-07-02
now ctx-matched to MTP/base (65536). Official nvidia/Qwen3.6-35B-A3B-NVFP4 on the AEON image, DFlash
n=11 via the small-page drafter, one-boot sweep. 0 errors.
- Result (conc 16): prefill 355.12 / decode 340.2 tok/s aggregate; 810/1000 prompts (hit the 600 s cap), 0 errors. vs MTP conc-16 (433.28): −21.5%.
- The cliff, essentially unchanged by the ctx fix. DFlash’s loss vs MTP steepens from the shallow ~−6 to −12% plateau through conc-2/4/8 to −21.5% at conc-16, on toward −24.8% at conc-32. As the batch fills, the ~7 wasted drafter forward passes per step compete directly with real decode work on this compute-bound MoE, and aggregate throughput falls further behind MTP the more streams share the GPU. The absolute decode number (340.2) is within 1.2% of the old ctx-40960 run’s 344.2 — at this batch size the KV-context difference barely mattered; the collapse is a genuine compute-bound effect, not a context artifact.
- Why — wasted draft compute. DFlash drafts 11 at ~26% acceptance / accept-len ~3.86 vs MTP’s 3 at ~67% / 3.0. Acceptance stays flat vs conc (workload-driven, matching every other point in this sweep within noise) — it’s the draft efficiency, not a drop in acceptance, that drives the collapse.
- Completes the ctx-matched DFlash money-chart line: decode 99.76 (c1) → 151.3 (c2) → 210.7 (c4) → 267.4 (c8) → 340.2 (c16) → 407.1 (c32); vs MTP that’s +0.7% → −6.1% → −9.3% → −12.0% → −21.5% → −24.8%. With context fully matched, DFlash’s headline “conc-1 win” is a wash, not a lead — the DFlash line never beats MTP at any concurrency on this checkpoint. Cliff still lands at c16→c32. Keep MTP for a mixed-workload gateway — now with no context caveat.
- One server lifetime for the whole six-point sweep → mem is the single ~111.3 GB reservation. TPOT 0.0 =
qwen3reasoning-parser client artifact. - Series:
c1(main) ·c2·c4·c8·c32. Matched MTP:-mtp-c16.