V100 SXM2 32GB vs P100 SXM2
NVIDIA V100 SXM2 32GB (Volta, 32 GB) against NVIDIA P100 SXM2 (Pascal, 16 GB): memory, compute, power and rental price, compared for LLM inference and training.
Pick two GPUs to compare
Side-by-Side Specifications
| Spec | V100 SXM2 32GB | P100 SXM2 |
|---|---|---|
| Architecture | Volta | Pascal |
| Memory | 32 GB HBM2 | 16 GB HBM2 |
| Memory bandwidth | 900 GB/s | 732 GB/s |
| FP16 tensor compute | 125 TFLOPS | 21.2 TFLOPS |
| INT8 tensor compute | 62.8 TOPS | 21.2 TOPS |
| Interconnect | NVLink 2.0 · 300 GB/s | NVLink 1.0 · 160 GB/s |
| TDP | 300 W | 300 W |
| Est. on-demand price | ~$2.00/h | ~$0.60/h |
| FP16 TFLOPS per $/h | 63 | 35 |
Highlighted values indicate the stronger spec. Hourly rates are indicative on-demand estimates.
Verdict
Raw performance: The V100 SXM2 32GB leads on FP16 tensor compute (5.9x advantage), which translates directly into higher token throughput for inference and shorter training steps.
Memory: With 32 GB per card, the V100 SXM2 32GB fits larger models on fewer GPUs — fewer cards means less inter-GPU communication and simpler deployments.
Value: At current on-demand rates, the V100 SXM2 32GB delivers more compute per dollar (63 vs 35 FP16 TFLOPS per $/h). If your model fits in its VRAM budget, it is usually the more economical choice.
GPUs Needed for Popular LLMs
Cards required to serve each model at 8-bit quantization (with 20% overhead for activations and KV cache).
| Model | VRAM (8-bit) | V100 SXM2 32GB | P100 SXM2 |
|---|---|---|---|
| GPT-5.6 Sol | 2682 GB | 84x | 168x |
| DeepSeek V4 Pro (671B) | 750 GB | 24x | 47x |
| Muse Spark 1.1 | 335 GB | 11x | 21x |
| Claude 5 Sonnet (175B) | 196 GB | 7x | 13x |
| Nova Premier (80B) | 89 GB | 3x | 6x |
| Nova Core (34B) | 38 GB | 2x | 3x |
| Nova Lite (12B) | 13 GB | 1x | 1x |
| Phi 3.5 (3.8B) | 4 GB | 1x | 1x |
Frequently Asked Questions
Which is better for LLM inference: V100 SXM2 32GB or P100 SXM2?
The V100 SXM2 32GB delivers more raw FP16 compute (125 TFLOPS) and the V100 SXM2 32GB offers the most memory per card (32 GB). For cost-efficiency, the V100 SXM2 32GB currently gives more FP16 TFLOPS per dollar of on-demand rental (63 vs 35 TFLOPS per $/h).
How much more memory does the V100 SXM2 32GB have?
The V100 SXM2 32GB has 32 GB of HBM2 versus 16 GB of HBM2 for the P100 SXM2 — a ratio of 2.00x in favor of the V100 SXM2 32GB. More VRAM per card means fewer GPUs to fit a given model.
Is the V100 SXM2 32GB or the P100 SXM2 cheaper to rent?
Estimated on-demand rates are ~$2.00/h for the V100 SXM2 32GB and ~$0.60/h for the P100 SXM2. Raw hourly price is only part of the story: normalize by throughput (TFLOPS per $/h) and by how many cards you need for your model's VRAM.
How do the V100 SXM2 32GB and P100 SXM2 compare on power?
The V100 SXM2 32GB has a TDP of 300W versus 300W for the P100 SXM2. FP16 compute per watt: 0.4 vs 0.1 TFLOPS/W.
Deploy on a GPU cloud
Rent the V100 SXM2 32GB or P100 SXM2 by the hour instead of buying hardware.