MiniMax-M2.7 · vLLM · NVFP4
status: blocked
Configuration
| Model | MiniMaxAI/MiniMax-M2.7 |
|---|---|
| Company | MiniMax |
| Family | MiniMax-M2 |
| Parameters | ~230B (≈10B active, MoE) |
| Engine | vLLM |
| Quant / precision | NVFP4 |
| Why this quant | saricles/MiniMax-M2.7-NVFP4-GB10 — an NVFP4 quant explicitly tagged for the GB10. Requested for queueing, but it is 130 GB and from an individual uploader — see status. |
| Download | saricles/MiniMax-M2.7-NVFP4-GB10 |
| Context window | 65536 |
| Input modalities | text |
Measured results
| Prefill tok/s | — |
|---|---|
| Decode tok/s | — |
| Peak memory (GB) | — |
Notes
BLOCKED — two reasons: it doesn’t fit, and the uploader isn’t trusted.
- Size: ~130 GB of weights > 121 GB unified ceiling. 27 safetensors shards (~5 GB each) sum to
130.6 GB (HF API, 2026-06-23). For an NVFP4 (W4) quant of a ~230B model this is on the high
side — likely keeps attention/embeddings/router in higher precision — but regardless it exceeds the
box. Unlike GGUF, a vLLM/SGLang NVFP4 load needs all weights resident and would OOM at load;
there is no paging fallback. The
-GB10tag notwithstanding, it does not fit a 121 GB GB10 with any KV pool. (If MiniMax shipped this for a larger-memory GB-class box, that’s not this machine.) - Trust:
sariclesis an individual uploader (1.0k downloads, 13 likes) — per the trusted-repo policy (model’s own org or a well-known quantizer only) this alone is a block until a trusted NVFP4 appears. No official MiniMax or ModelOpt NVFP4 of M2.7 was found.
No vLLM-viable NVFP4 path for MiniMax-M2.7 on this box right now — both the size and the trust gates fail. Revisit if (a) an official/ModelOpt NVFP4 under ~115 GB is published, or (b) the user explicitly authorizes the saricles repo and it can be confirmed to fit (it currently cannot).
Sources: saricles/MiniMax-M2.7-NVFP4-GB10.