DeepSeek AI

DeepSeek V3.1 (32B)

Dense model optimized for quality and latency balance

Model Summary

Family

DeepSeek

Version

3.1

Parameters

32B (est.)

Parameter counts for closed models are estimates; vendors rarely publish exact sizes.

VRAM Requirements by Quantization

Memory needed to serve DeepSeek V3.1 (32B) for inference, including a 20% overhead for activations and KV cache.

PrecisionVRAM neededSmallest single GPU that fits
INT4 (4-bit)17.88 GBNVIDIA L4 (24 GB)
INT8 (8-bit)35.76 GBNVIDIA L40S (48 GB)
FP16 (16-bit)71.53 GBNVIDIA H100 SXM5 (80 GB)
FP32 (32-bit)143.05 GBNVIDIA B200 SXM (192 GB)

Recommended GPU Configurations

Cheapest on-demand configurations to serve DeepSeek V3.1 (32B) at 8-bit (36 GB VRAM).

3x NVIDIA T4

48 GB total VRAM · Turing

~$1.50/h

3x NVIDIA P100 SXM2

48 GB total VRAM · Pascal

~$1.80/h

1x NVIDIA A40

48 GB total VRAM · Ampere

~$1.80/h

Quick GPU Planning

Use the calculator pre-filled with this exact version to estimate memory, speed, and compute requirements in a few clicks.

Access Pre-filled Calculator

Frequently Asked Questions

How much VRAM do you need to run DeepSeek V3.1 (32B)?

With an estimated 32B parameters, DeepSeek V3.1 (32B) needs roughly 36 GB of VRAM in 8-bit (INT8), 18 GB in 4-bit, and 72 GB in FP16, including a 20% overhead for activations and KV cache.

Which GPUs can run DeepSeek V3.1 (32B)?

At 8-bit quantization, the most cost-effective option is 3x NVIDIA T4 (48 GB combined VRAM, around $1.50/hour on-demand). Higher-end cards like the NVIDIA B200 or AMD MI355X reduce the GPU count needed.

Can DeepSeek V3.1 (32B) run on a single GPU?

Yes. In 8-bit, a single NVIDIA L40S (48 GB) fits the model.

How much does it cost to serve DeepSeek V3.1 (32B) in the cloud?

Renting 3x T4 costs on the order of $1.50/hour, i.e. about $1,095/month running 24/7. Actual prices vary by provider and commitment; spot and reserved capacity can be significantly cheaper.