GPU Specifications / NVIDIA

NVIDIA V100 SXM2 32GB

Volta datacenter GPU with 32 GB of HBM2 memory, 900 GB/s of bandwidth and up to 125 TFLOPS of FP16 tensor compute.

Key Specifications

Memory

32 GB HBM2

Memory Bandwidth

900 GB/s

TDP

300 W

Architecture

Volta

Interconnect

NVLink 2.0 · 300 GB/s

Est. On-demand Price

~$2.00/h

Hourly rates are indicative on-demand estimates; actual pricing varies by provider and commitment.

Compute Performance

PrecisionPeak throughput
FP64 (double precision)7.8 TFLOPS
FP32 (single precision)15.7 TFLOPS
FP2431.4 TFLOPS
FP16 (tensor)125 TFLOPS
INT8 (tensor)62.8 TFLOPS
INT4 (tensor)62.8 TFLOPS

Tensor figures use the vendor's peak numbers (with structured sparsity where supported).

System Requirements

Recommended CPU

Intel Xeon Gold 6148 or AMD EPYC 7601

Max VRAM per node (8 GPUs)

256 GB

System RAM (min / recommended)

192 / 384 GB

Minimum PSU

1000 W

LLMs on the V100 SXM2 32GB

Number of V100 SXM2 32GB GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).

ModelParamsVRAM (8-bit)GPUs needed
GPT-5.6 Sol2400B2682 GB84x V100 SXM2 32GB
GPT-5 Flagship2100B2347 GB74x V100 SXM2 32GB
GPT-5.6 Luna1400B1565 GB49x V100 SXM2 32GB
Kimi K3 (1.2T)1200B1341 GB42x V100 SXM2 32GB
Kimi K2.6 (1T)1000B1118 GB35x V100 SXM2 32GB
GPT-5.6 Terra800B894 GB28x V100 SXM2 32GB
DeepSeek V4 Pro (671B)671B750 GB24x V100 SXM2 32GB
Llama 4 Behemoth (500B)500B559 GB18x V100 SXM2 32GB
Claude 5 Fable (480B)480B536 GB17x V100 SXM2 32GB
GLM 5.2 (400B)400B447 GB14x V100 SXM2 32GB
Grok 4.5350B391 GB13x V100 SXM2 32GB
Claude 4.8 Opus (300B)300B335 GB11x V100 SXM2 32GB
Muse Spark 1.1300B335 GB11x V100 SXM2 32GB
Grok 4270B302 GB10x V100 SXM2 32GB
Gemini 3.1 Pro250B279 GB9x V100 SXM2 32GB
Qwen 3.7 Max (235B)235B263 GB9x V100 SXM2 32GB
Mistral Large 3 (200B)200B224 GB7x V100 SXM2 32GB
Grok 3 Mini190B212 GB7x V100 SXM2 32GB
Claude 5 Sonnet (175B)175B196 GB7x V100 SXM2 32GB
Gemini 3.5 Flash150B168 GB6x V100 SXM2 32GB
Gemini 2.5 Flash140B156 GB5x V100 SXM2 32GB
Llama 4 Maverick (128B)128B143 GB5x V100 SXM2 32GB
DeepSeek V4 Flash (120B)120B134 GB5x V100 SXM2 32GB
Qwen 3.6 Plus (110B)110B123 GB4x V100 SXM2 32GB
Nova Premier (80B)80B89 GB3x V100 SXM2 32GB
Qwen 3 Coder-Next (80B)80B89 GB3x V100 SXM2 32GB
Claude 4.5 Haiku (70B)70B78 GB3x V100 SXM2 32GB
Llama 3.3 Instruct (70B)70B78 GB3x V100 SXM2 32GB
Mistral Medium 3.5 (70B)70B78 GB3x V100 SXM2 32GB
Yi 1.5 (40B)40B45 GB2x V100 SXM2 32GB
Nova Core (34B)34B38 GB2x V100 SXM2 32GB
DeepSeek V3.1 (32B)32B36 GB2x V100 SXM2 32GB
Gemma 3 (27B)27B30 GB1x V100 SXM2 32GB
Mistral Small 4 (24B)24B27 GB1x V100 SXM2 32GB
Yi 1.5 (15B)15B17 GB1x V100 SXM2 32GB
Phi 4 (14B)14B16 GB1x V100 SXM2 32GB
Nova Lite (12B)12B13 GB1x V100 SXM2 32GB
Llama 3.2 Instruct (11B)11B12 GB1x V100 SXM2 32GB
Gemma 3 (9B)9B10 GB1x V100 SXM2 32GB
Yi 1.5 Lite (9B)9B10 GB1x V100 SXM2 32GB
Phi 4 Mini (7B)7B8 GB1x V100 SXM2 32GB
Phi 3.5 (3.8B)3.8B4 GB1x V100 SXM2 32GB

Frequently Asked Questions

How much VRAM does the NVIDIA V100 SXM2 32GB have?

The NVIDIA V100 SXM2 32GB has 32 GB of HBM2 memory with 900 GB/s of memory bandwidth.

Which LLMs can run on a single V100 SXM2 32GB?

At 8-bit quantization, a single V100 SXM2 32GB (32 GB) can serve models up to roughly 27B parameters, such as Gemma 3 (27B). Larger models require multiple GPUs or more aggressive quantization.

How much does it cost to rent a NVIDIA V100 SXM2 32GB?

On-demand cloud pricing for the V100 SXM2 32GB is around $2.00/hour, i.e. about $1,460/month running 24/7. Actual prices vary by provider, region, and commitment.

What are the power and system requirements of the NVIDIA V100 SXM2 32GB?

The V100 SXM2 32GB has a TDP of 300W. A power supply of at least 1000W per GPU is recommended. Recommended host CPUs: Intel Xeon Gold 6148 or AMD EPYC 7601.