GPU Specifications / NVIDIA

NVIDIA P100 SXM2

Pascal datacenter GPU with 16 GB of HBM2 memory, 732 GB/s of bandwidth and up to 21.2 TFLOPS of FP16 tensor compute.

Key Specifications

Memory

16 GB HBM2

Memory Bandwidth

732 GB/s

TDP

300 W

Architecture

Pascal

Interconnect

NVLink 1.0 · 160 GB/s

Est. On-demand Price

~$0.60/h

Hourly rates are indicative on-demand estimates; actual pricing varies by provider and commitment.

Compute Performance

PrecisionPeak throughput
FP64 (double precision)5.3 TFLOPS
FP32 (single precision)10.6 TFLOPS
FP2421.2 TFLOPS
FP16 (tensor)21.2 TFLOPS
INT8 (tensor)21.2 TFLOPS
INT4 (tensor)21.2 TFLOPS

Tensor figures use the vendor's peak numbers (with structured sparsity where supported).

System Requirements

Recommended CPU

Intel Xeon E5-2698 v4

Max VRAM per node (8 GPUs)

128 GB

System RAM (min / recommended)

128 / 256 GB

Minimum PSU

800 W

LLMs on the P100 SXM2

Number of P100 SXM2 GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).

ModelParamsVRAM (8-bit)GPUs needed
GPT-5.6 Sol2400B2682 GB168x P100 SXM2
GPT-5 Flagship2100B2347 GB147x P100 SXM2
GPT-5.6 Luna1400B1565 GB98x P100 SXM2
Kimi K3 (1.2T)1200B1341 GB84x P100 SXM2
Kimi K2.6 (1T)1000B1118 GB70x P100 SXM2
GPT-5.6 Terra800B894 GB56x P100 SXM2
DeepSeek V4 Pro (671B)671B750 GB47x P100 SXM2
Llama 4 Behemoth (500B)500B559 GB35x P100 SXM2
Claude 5 Fable (480B)480B536 GB34x P100 SXM2
GLM 5.2 (400B)400B447 GB28x P100 SXM2
Grok 4.5350B391 GB25x P100 SXM2
Claude 4.8 Opus (300B)300B335 GB21x P100 SXM2
Muse Spark 1.1300B335 GB21x P100 SXM2
Grok 4270B302 GB19x P100 SXM2
Gemini 3.1 Pro250B279 GB18x P100 SXM2
Qwen 3.7 Max (235B)235B263 GB17x P100 SXM2
Mistral Large 3 (200B)200B224 GB14x P100 SXM2
Grok 3 Mini190B212 GB14x P100 SXM2
Claude 5 Sonnet (175B)175B196 GB13x P100 SXM2
Gemini 3.5 Flash150B168 GB11x P100 SXM2
Gemini 2.5 Flash140B156 GB10x P100 SXM2
Llama 4 Maverick (128B)128B143 GB9x P100 SXM2
DeepSeek V4 Flash (120B)120B134 GB9x P100 SXM2
Qwen 3.6 Plus (110B)110B123 GB8x P100 SXM2
Nova Premier (80B)80B89 GB6x P100 SXM2
Qwen 3 Coder-Next (80B)80B89 GB6x P100 SXM2
Claude 4.5 Haiku (70B)70B78 GB5x P100 SXM2
Llama 3.3 Instruct (70B)70B78 GB5x P100 SXM2
Mistral Medium 3.5 (70B)70B78 GB5x P100 SXM2
Yi 1.5 (40B)40B45 GB3x P100 SXM2
Nova Core (34B)34B38 GB3x P100 SXM2
DeepSeek V3.1 (32B)32B36 GB3x P100 SXM2
Gemma 3 (27B)27B30 GB2x P100 SXM2
Mistral Small 4 (24B)24B27 GB2x P100 SXM2
Yi 1.5 (15B)15B17 GB2x P100 SXM2
Phi 4 (14B)14B16 GB1x P100 SXM2
Nova Lite (12B)12B13 GB1x P100 SXM2
Llama 3.2 Instruct (11B)11B12 GB1x P100 SXM2
Gemma 3 (9B)9B10 GB1x P100 SXM2
Yi 1.5 Lite (9B)9B10 GB1x P100 SXM2
Phi 4 Mini (7B)7B8 GB1x P100 SXM2
Phi 3.5 (3.8B)3.8B4 GB1x P100 SXM2

Frequently Asked Questions

How much VRAM does the NVIDIA P100 SXM2 have?

The NVIDIA P100 SXM2 has 16 GB of HBM2 memory with 732 GB/s of memory bandwidth.

Which LLMs can run on a single P100 SXM2?

At 8-bit quantization, a single P100 SXM2 (16 GB) can serve models up to roughly 14B parameters, such as Phi 4 (14B). Larger models require multiple GPUs or more aggressive quantization.

How much does it cost to rent a NVIDIA P100 SXM2?

On-demand cloud pricing for the P100 SXM2 is around $0.60/hour, i.e. about $438/month running 24/7. Actual prices vary by provider, region, and commitment.

What are the power and system requirements of the NVIDIA P100 SXM2?

The P100 SXM2 has a TDP of 300W. A power supply of at least 800W per GPU is recommended. Recommended host CPUs: Intel Xeon E5-2698 v4.