Browse by tag
Every configuration is tagged in seven kinds, plus three provenance tags for models we compressed ourselves. Each has its own page, where every tag value lists its configurations in a table like the homepage. Pick a kind:
- Model — a per-model slug (e.g.
gemma-4-31b). Every run of one model shares it, so it groups all of a model’s engines, quants, and speculative variants. - Lab — the lab that released the model (e.g. NVIDIA, OpenAI).
- Family — model family (e.g. Nemotron, Llama, Gemma).
- Quant / precision — weight format (e.g. NVFP4, FP8, Q4_K_M).
- Size bucket — by total params (
≤4B,5-15B,16-40B,41-130B,130B+). - Concurrency — parallel serves per run (e.g.
conc-32); the benchmark’s fixed-load axis. - Spark recipe — models with native DGX Spark support.
- REAP — base checkpoint pruned with REAP (router-weighted expert-activation pruning).
- Self-REAP — models we REAP-pruned ourselves; the page carries the how-to (none pruned yet).
- Self-Quantized — models we quantized ourselves to NVFP4; the page carries the streaming quantizer + full recipe.
Model Lab Family Quant Size Concurrency Spark recipe REAP Self-REAP Self-Quantized