GLM-4.7-Flash · vLLM · FP8
status: blocked
Configuration
| Model | zai-org/GLM-4.7-Flash |
|---|---|
| Company | Zhipu AI |
| Family | GLM |
| Parameters | 31B (MoE) |
| Engine | vLLM |
| Quant / precision | FP8 |
| Why this quant | Near-BF16 quality at half the bytes; official FP8 weights published. |
| Download | zai-org/GLM-4.7-Flash |
| Context window | 131072 |
| Input modalities | text |
Measured results
| Prefill tok/s | — |
|---|---|
| Decode tok/s | — |
| Peak memory (GB) | — |
Full run command
# blocked — no confirmable repo (see Notes)
Notes
Blocked — the named model does not resolve to a real HF repo, and no clean substitute fits the intended slot. Verified 2026-06-22.
zai-org/GLM-4.7-Flash(andglm-4-7-flash,GLM-4-Flash) → 404 on HF. The stub describes a ~31B MoE GLM “Flash”, which does not exist yet.- The nearby real GLMs don’t match:
zai-org/GLM-4.5-Airis a 106B/12B MoE (belongs in the 41-130B / 130B+ buckets, not 16-40B), andzai-org/GLM-4-32B-0414is a 32B dense model, not a Flash MoE. Substituting either would change the model class this config is meant to capture. - Per the “when unsure, BLOCK — don’t guess” policy, left for human review rather than silently
swapped. If the intent is the dense 32B, point
source_repoatzai-org/GLM-4-32B-0414; if it’s the small MoE, usezai-org/GLM-4.5-Air(and re-bucket by size).