Model planner
Can I run it?
Pick your hardware, a quantization and a context length. See what fits in memory and the decode ceiling your memory bandwidth allows.
Weights
Most popular, about 4.85 bits per weight.
Context
- Memory
- 12 GB
- Published peak
- 504 GB/s
10 of 22 models run on this machine
- Llama 3.2 1BFits1.8 of 12 GB670tok/s
- Llama 3.2 3BFits3.2 of 12 GB259tok/s
- Llama 3.1 8BFits6.4 of 12 GB104tok/s
- Llama 3.3 70BToo large45.9 of 12 GB—
- Llama 3.1 405BToo large252.5 of 12 GB—
- Phi-3.5 miniFits3.6 of 12 GB218tok/s
- Phi-4Fits10.7 of 12 GB56.6tok/s
- Gemma 2 9BFits7.2 of 12 GB90.0tok/s
- Gemma 2 27BToo large18.7 of 12 GB—
- Mistral 7BFits5.9 of 12 GB115tok/s
- Mistral Small 24BToo large16.4 of 12 GB—
- Mixtral 8x7B12.9B activeToo large31.0 of 12 GB—
- Qwen 2.5 7BFits6.2 of 12 GB109tok/s
- Qwen 2.5 14BFits10.8 of 12 GB56.2tok/s
- Qwen 2.5 32BToo large22.2 of 12 GB—
- Qwen 2.5 Coder 32BToo large22.2 of 12 GB—
- Qwen 2.5 72BToo large47.2 of 12 GB—
- DeepSeek-R1 distill 8BFits6.4 of 12 GB104tok/s
- DeepSeek-R1 distill 32BToo large22.2 of 12 GB—
- DeepSeek-R137B activeToo large414.8 of 12 GB—
- gpt-oss-20b3.6B activeToo large14.7 of 12 GB—
- gpt-oss-120b5.1B activeToo large74.6 of 12 GB—
The ceiling is memory bandwidth divided by the bytes of weights read per token. It is an upper bound from the memory side; real runtimes land below it. Memory need is the weights plus an estimate for the KV cache and runtime.
Published figures are the ceiling. Yours may differ.
Run the benchmark, then enter your measured read throughput above to plan with what your machine really delivers.