Qwen3.6-27B AEON Uncensored · vLLM-ultimate (custom) · NVFP4 + DFlash · conc 8
status: done
Configuration
| Model | AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored |
|---|---|
| Company | Alibaba |
| Family | Qwen |
| Parameters | 27B (dense) + DFlash external drafter |
| Engine | vLLM (aeon-vllm-ultimate custom container) + DFlash (z-lab drafter, num_speculative_tokens 12) (speculative decoding) |
| Quant / precision | NVFP4 (XS mixed-precision) |
| Why this quant | Same AEON NVFP4-XS model served via the card's "DGX Spark production" recipe — custom container ghcr.io/aeon-7/aeon-vllm-ultimate:latest (vLLM 0.23.0) + external z-lab/Qwen3.6-27B-DFlash drafter, DFlash num_speculative_tokens=12. User's explicit request. SAFETY — untrusted image; NO creds, models READ-ONLY. |
| Download | AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS |
| Context window | 65536 |
| Input modalities | text, image, video (served text-only here) |
Measured results
| Prefill tok/s | 124.32 |
|---|---|
| Decode tok/s | 127.14 |
| Peak memory (GB) | 106.17 (system MemAvailable delta (10s sampling) — custom vLLM static KV reservation (util 0.85)) |
| Completed | 2026-06-24 01:11 +0800 |
Full run command
# UNTRUSTED third-party container — NO creds, full cached repo root + drafter mounted READ-ONLY.
scripts/bench-aeon-ultimate-serving.sh 65536 8 1000 900 256
# 456 prompts completed before the 900 s cap (0 errors). TTFT median 14875.8 ms.
Notes
Custom AEON-ultimate container + DFlash, conc 8. Part of the 1/8/32 sweep on the card’s “DGX Spark production” recipe (custom vLLM 0.23.0 + external z-lab DFlash drafter). Mid-point between single-stream and the c=32 throughput point. See the conc-1 page for the safety posture (untrusted image, no creds, models read-only) and drafter details.
- Result (conc 8): prefill 124.32 tok/s, decode 127.14 tok/s aggregate; 456 prompts
completed before the 900 s cap, with 0 errors. Median TTFT 14875.8 ms. So batching helps
aggregate throughput relative to
conc=1, but the path still misses the benchmark target badly. - Acceptance: DFlash again showed a front-loaded 12-token draft profile with mean acceptance
length ~3.3-4.3 and avg draft acceptance ~19-28%. That is consistent with the
conc=1page: some speculative work is paying off, but not nearly enough to justify the observed latency and cap behavior. - Takeaway: the custom AEON container is not rescuing throughput under moderate batch either. It remains a stable but slow serving path on this workload.