Meta
Llama 4 Maverick (128B)
Balanced large model for quality and serving efficiency
Model Summary
Family
Llama 4
Version
4.1
Parameters
128B (est.)
Parameter counts for closed models are estimates; vendors rarely publish exact sizes.
VRAM Requirements by Quantization
Memory needed to serve Llama 4 Maverick (128B) for inference, including a 20% overhead for activations and KV cache.
| Precision | VRAM needed | Smallest single GPU that fits |
|---|---|---|
| INT4 (4-bit) | 71.53 GB | NVIDIA H100 SXM5 (80 GB) |
| INT8 (8-bit) | 143.05 GB | NVIDIA B200 SXM (192 GB) |
| FP16 (16-bit) | 286.10 GB | NVIDIA B300 SXM (288 GB) |
| FP32 (32-bit) | 572.20 GB | Multi-GPU required |
Recommended GPU Configurations
Cheapest on-demand configurations to serve Llama 4 Maverick (128B) at 8-bit (143 GB VRAM).
9x NVIDIA T4
144 GB total VRAM · Turing · multi-node
~$4.50/h
2x AMD Instinct MI250X
256 GB total VRAM · CDNA 2
~$5.00/h
9x NVIDIA P100 SXM2
144 GB total VRAM · Pascal · multi-node
~$5.40/h
Quick GPU Planning
Use the calculator pre-filled with this exact version to estimate memory, speed, and compute requirements in a few clicks.
Access Pre-filled CalculatorFrequently Asked Questions
How much VRAM do you need to run Llama 4 Maverick (128B)?
With an estimated 128B parameters, Llama 4 Maverick (128B) needs roughly 143 GB of VRAM in 8-bit (INT8), 72 GB in 4-bit, and 286 GB in FP16, including a 20% overhead for activations and KV cache.
Which GPUs can run Llama 4 Maverick (128B)?
At 8-bit quantization, the most cost-effective option is 9x NVIDIA T4 (144 GB combined VRAM, around $4.50/hour on-demand). Higher-end cards like the NVIDIA B200 or AMD MI355X reduce the GPU count needed.
Can Llama 4 Maverick (128B) run on a single GPU?
Yes. In 8-bit, a single NVIDIA B200 SXM (192 GB) fits the model.
How much does it cost to serve Llama 4 Maverick (128B) in the cloud?
Renting 9x T4 costs on the order of $4.50/hour, i.e. about $3,285/month running 24/7. Actual prices vary by provider and commitment; spot and reserved capacity can be significantly cheaper.
Deploy on a GPU cloud
Rent 9x T4 by the hour instead of buying hardware.