Specs Comparisons

V100 SXM2 32GB vs P100 SXM2

NVIDIA V100 SXM2 32GB (Volta, 32 GB) against NVIDIA P100 SXM2 (Pascal, 16 GB): memory, compute, power and rental price, compared for LLM inference and training.

Pick two GPUs to compare

vs

Side-by-Side Specifications

SpecV100 SXM2 32GBP100 SXM2
ArchitectureVoltaPascal
Memory32 GB HBM216 GB HBM2
Memory bandwidth900 GB/s732 GB/s
FP16 tensor compute125 TFLOPS21.2 TFLOPS
INT8 tensor compute62.8 TOPS21.2 TOPS
InterconnectNVLink 2.0 · 300 GB/sNVLink 1.0 · 160 GB/s
TDP300 W300 W
Est. on-demand price~$2.00/h~$0.60/h
FP16 TFLOPS per $/h6335

Highlighted values indicate the stronger spec. Hourly rates are indicative on-demand estimates.

Verdict

Raw performance: The V100 SXM2 32GB leads on FP16 tensor compute (5.9x advantage), which translates directly into higher token throughput for inference and shorter training steps.

Memory: With 32 GB per card, the V100 SXM2 32GB fits larger models on fewer GPUs — fewer cards means less inter-GPU communication and simpler deployments.

Value: At current on-demand rates, the V100 SXM2 32GB delivers more compute per dollar (63 vs 35 FP16 TFLOPS per $/h). If your model fits in its VRAM budget, it is usually the more economical choice.

GPUs Needed for Popular LLMs

Cards required to serve each model at 8-bit quantization (with 20% overhead for activations and KV cache).

ModelVRAM (8-bit)V100 SXM2 32GBP100 SXM2
GPT-5.6 Sol2682 GB84x168x
DeepSeek V4 Pro (671B)750 GB24x47x
Muse Spark 1.1335 GB11x21x
Claude 5 Sonnet (175B)196 GB7x13x
Nova Premier (80B)89 GB3x6x
Nova Core (34B)38 GB2x3x
Nova Lite (12B)13 GB1x1x
Phi 3.5 (3.8B)4 GB1x1x

Frequently Asked Questions

Which is better for LLM inference: V100 SXM2 32GB or P100 SXM2?

The V100 SXM2 32GB delivers more raw FP16 compute (125 TFLOPS) and the V100 SXM2 32GB offers the most memory per card (32 GB). For cost-efficiency, the V100 SXM2 32GB currently gives more FP16 TFLOPS per dollar of on-demand rental (63 vs 35 TFLOPS per $/h).

How much more memory does the V100 SXM2 32GB have?

The V100 SXM2 32GB has 32 GB of HBM2 versus 16 GB of HBM2 for the P100 SXM2 — a ratio of 2.00x in favor of the V100 SXM2 32GB. More VRAM per card means fewer GPUs to fit a given model.

Is the V100 SXM2 32GB or the P100 SXM2 cheaper to rent?

Estimated on-demand rates are ~$2.00/h for the V100 SXM2 32GB and ~$0.60/h for the P100 SXM2. Raw hourly price is only part of the story: normalize by throughput (TFLOPS per $/h) and by how many cards you need for your model's VRAM.

How do the V100 SXM2 32GB and P100 SXM2 compare on power?

The V100 SXM2 32GB has a TDP of 300W versus 300W for the P100 SXM2. FP16 compute per watt: 0.4 vs 0.1 TFLOPS/W.