Google

Gemma 3 (27B)

Google open model with strong multilingual capabilities

Model Summary

Family

Gemma

Version

3.0

Parameters

27B (est.)

Parameter counts for closed models are estimates; vendors rarely publish exact sizes.

VRAM Requirements by Quantization

Memory needed to serve Gemma 3 (27B) for inference, including a 20% overhead for activations and KV cache.

PrecisionVRAM neededSmallest single GPU that fits
INT4 (4-bit)15.09 GBNVIDIA P100 SXM2 (16 GB)
INT8 (8-bit)30.17 GBNVIDIA V100 SXM2 32GB (32 GB)
FP16 (16-bit)60.35 GBNVIDIA H100 SXM5 (80 GB)
FP32 (32-bit)120.70 GBAMD Instinct MI250X (128 GB)

Recommended GPU Configurations

Cheapest on-demand configurations to serve Gemma 3 (27B) at 8-bit (30 GB VRAM).

2x NVIDIA T4

32 GB total VRAM · Turing

~$1.00/h

2x NVIDIA P100 SXM2

32 GB total VRAM · Pascal

~$1.20/h

1x NVIDIA A40

48 GB total VRAM · Ampere

~$1.80/h

Quick GPU Planning

Use the calculator pre-filled with this exact version to estimate memory, speed, and compute requirements in a few clicks.

Access Pre-filled Calculator

Frequently Asked Questions

How much VRAM do you need to run Gemma 3 (27B)?

With an estimated 27B parameters, Gemma 3 (27B) needs roughly 30 GB of VRAM in 8-bit (INT8), 15 GB in 4-bit, and 60 GB in FP16, including a 20% overhead for activations and KV cache.

Which GPUs can run Gemma 3 (27B)?

At 8-bit quantization, the most cost-effective option is 2x NVIDIA T4 (32 GB combined VRAM, around $1.00/hour on-demand). Higher-end cards like the NVIDIA B200 or AMD MI355X reduce the GPU count needed.

Can Gemma 3 (27B) run on a single GPU?

Yes. In 8-bit, a single NVIDIA V100 SXM2 32GB (32 GB) fits the model.

How much does it cost to serve Gemma 3 (27B) in the cloud?

Renting 2x T4 costs on the order of $1.00/hour, i.e. about $730/month running 24/7. Actual prices vary by provider and commitment; spot and reserved capacity can be significantly cheaper.

Other Gemma Versions