Qwen3.6-27B AEON Uncensored · vLLM-ultimate (custom) · NVFP4 + DFlash · conc 1
status: done
Configuration
| Model | AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored |
|---|---|
| Company | Alibaba |
| Family | Qwen |
| Parameters | 27B (dense) + DFlash external drafter |
| Engine | vLLM (aeon-vllm-ultimate custom container) + DFlash (z-lab drafter, num_speculative_tokens 12) (speculative decoding) |
| Quant / precision | NVFP4 (XS mixed-precision) |
| Why this quant | Same AEON NVFP4-XS model as the native-MTP configs, but served via the card's "DGX Spark production" recipe — the custom third-party container ghcr.io/aeon-7/aeon-vllm-ultimate:latest (vLLM 0.23.0) + the external z-lab/Qwen3.6-27B-DFlash drafter mounted at /drafter, DFlash spec-decode num_speculative_tokens=12. Run at the user's explicit request, reversing the earlier "untrusted container declined" call. SAFETY — image is untrusted; run with NO credentials (HF_TOKEN withheld), both models mounted READ-ONLY. |
| Download | AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS |
| Context window | 65536 |
| Input modalities | text, image, video (served text-only here) |
Measured results
| Prefill tok/s | 38.69 |
|---|---|
| Decode tok/s | 29.87 |
| Peak memory (GB) | 106.96 (system MemAvailable delta (10s sampling) — custom vLLM static KV reservation (util 0.85)) |
| Completed | 2026-06-24 00:47 +0800 |
Full run command
# UNTRUSTED third-party container — NO creds, model repo root + drafter mounted READ-ONLY.
# The cached HF snapshot itself contains symlinks into ../blobs, so mount the full repo root,
# not only the snapshot directory.
scripts/bench-aeon-ultimate-serving.sh 65536 1 1000 900 256
# 106 prompts completed before the 900 s cap (0 errors). TTFT median 8111.8 ms.
Notes
DONE — the custom AEON DFlash path is dramatically slower than expected even at conc-1.
- Result (conc 1): prefill 38.69 tok/s, decode 29.87 tok/s aggregate; only 106 prompts completed before the 900 s cap, with 0 errors. Median TTFT 8111.8 ms. This is not even close to the card’s claimed DGX Spark production behavior.
- This is a real performance failure, not a launch failure: after fixing the local mount shape
(the cached snapshot contains symlinks into Hugging Face
blobs/, so the full repo root must be mounted), the custom image loaded the base asQwen3_5ForConditionalGenerationand the drafter asDFlashDraftModelon its ownv0.23.0+aeon.sm121a.dflashfork. The benchmark then ran stably to the time cap with no request errors. - Acceptance was not the main blocker: DFlash logged mean acceptance length ~3.0-5.8, but the avg draft acceptance rate only sat around ~18-40% across windows because the 12-token draft has a very front-loaded per-position survival curve. The first few draft positions accept well, later ones collapse quickly. That is enough to explain some underperformance, but not the sheer magnitude of this slowdown.
- Bottom line: on this box and workload, the AEON custom container + external DFlash drafter is not competitive with the already-done stock-vLLM native-MTP path. This single-stream point is bad enough that the higher-concurrency points are likely measuring how quickly it degrades from an already poor baseline, not uncovering a hidden win.