01.AI
Yi 1.5 Lite (9B)
Compact Yi variant focused on efficiency and low cost
Model Summary
Family
Yi
Version
1.5
Parameters
9B (est.)
Parameter counts for closed models are estimates; vendors rarely publish exact sizes.
VRAM Requirements by Quantization
Memory needed to serve Yi 1.5 Lite (9B) for inference, including a 20% overhead for activations and KV cache.
| Precision | VRAM needed | Smallest single GPU that fits |
|---|---|---|
| INT4 (4-bit) | 5.03 GB | NVIDIA P100 SXM2 (16 GB) |
| INT8 (8-bit) | 10.06 GB | NVIDIA P100 SXM2 (16 GB) |
| FP16 (16-bit) | 20.12 GB | NVIDIA L4 (24 GB) |
| FP32 (32-bit) | 40.23 GB | NVIDIA L40S (48 GB) |
Recommended GPU Configurations
Cheapest on-demand configurations to serve Yi 1.5 Lite (9B) at 8-bit (10 GB VRAM).
1x NVIDIA T4
16 GB total VRAM · Turing
~$0.50/h
1x NVIDIA P100 SXM2
16 GB total VRAM · Pascal
~$0.60/h
1x NVIDIA L4
24 GB total VRAM · Ada Lovelace
~$1.00/h
Quick GPU Planning
Use the calculator pre-filled with this exact version to estimate memory, speed, and compute requirements in a few clicks.
Access Pre-filled CalculatorFrequently Asked Questions
How much VRAM do you need to run Yi 1.5 Lite (9B)?
With an estimated 9B parameters, Yi 1.5 Lite (9B) needs roughly 10 GB of VRAM in 8-bit (INT8), 5 GB in 4-bit, and 20 GB in FP16, including a 20% overhead for activations and KV cache.
Which GPUs can run Yi 1.5 Lite (9B)?
At 8-bit quantization, the most cost-effective option is 1x NVIDIA T4 (16 GB combined VRAM, around $0.50/hour on-demand). Higher-end cards like the NVIDIA B200 or AMD MI355X reduce the GPU count needed.
Can Yi 1.5 Lite (9B) run on a single GPU?
Yes. In 8-bit, a single NVIDIA P100 SXM2 (16 GB) fits the model.
How much does it cost to serve Yi 1.5 Lite (9B) in the cloud?
Renting 1x T4 costs on the order of $0.50/hour, i.e. about $365/month running 24/7. Actual prices vary by provider and commitment; spot and reserved capacity can be significantly cheaper.
Deploy on a GPU cloud
Rent 1x T4 by the hour instead of buying hardware.