01.AI
Yi 1.5 (40B)
Updated Yi flagship model with stronger multilingual quality
Model Summary
Family
Yi
Version
1.5
Parameters
40B (est.)
Parameter counts for closed models are estimates; vendors rarely publish exact sizes.
VRAM Requirements by Quantization
Memory needed to serve Yi 1.5 (40B) for inference, including a 20% overhead for activations and KV cache.
| Precision | VRAM needed | Smallest single GPU that fits |
|---|---|---|
| INT4 (4-bit) | 22.35 GB | NVIDIA L4 (24 GB) |
| INT8 (8-bit) | 44.70 GB | NVIDIA L40S (48 GB) |
| FP16 (16-bit) | 89.41 GB | NVIDIA RTX PRO 6000 Blackwell (96 GB) |
| FP32 (32-bit) | 178.81 GB | NVIDIA B200 SXM (192 GB) |
Recommended GPU Configurations
Cheapest on-demand configurations to serve Yi 1.5 (40B) at 8-bit (45 GB VRAM).
3x NVIDIA T4
48 GB total VRAM · Turing
~$1.50/h
3x NVIDIA P100 SXM2
48 GB total VRAM · Pascal
~$1.80/h
1x NVIDIA A40
48 GB total VRAM · Ampere
~$1.80/h
Quick GPU Planning
Use the calculator pre-filled with this exact version to estimate memory, speed, and compute requirements in a few clicks.
Access Pre-filled CalculatorFrequently Asked Questions
How much VRAM do you need to run Yi 1.5 (40B)?
With an estimated 40B parameters, Yi 1.5 (40B) needs roughly 45 GB of VRAM in 8-bit (INT8), 22 GB in 4-bit, and 89 GB in FP16, including a 20% overhead for activations and KV cache.
Which GPUs can run Yi 1.5 (40B)?
At 8-bit quantization, the most cost-effective option is 3x NVIDIA T4 (48 GB combined VRAM, around $1.50/hour on-demand). Higher-end cards like the NVIDIA B200 or AMD MI355X reduce the GPU count needed.
Can Yi 1.5 (40B) run on a single GPU?
Yes. In 8-bit, a single NVIDIA L40S (48 GB) fits the model.
How much does it cost to serve Yi 1.5 (40B) in the cloud?
Renting 3x T4 costs on the order of $1.50/hour, i.e. about $1,095/month running 24/7. Actual prices vary by provider and commitment; spot and reserved capacity can be significantly cheaper.
Deploy on a GPU cloud
Rent 3x T4 by the hour instead of buying hardware.