GPU Specifications / NVIDIA

NVIDIA H200 SXM5

Hopper datacenter GPU with 141 GB of HBM3e memory, 4800 GB/s of bandwidth and up to 1,979 TFLOPS of FP16 tensor compute.

Key Specifications

Memory

141 GB HBM3e

Memory Bandwidth

4,800 GB/s

TDP

700 W

Architecture

Hopper

Interconnect

NVLink 4.0 · 900 GB/s

Est. On-demand Price

~$10.00/h

Hourly rates are indicative on-demand estimates; actual pricing varies by provider and commitment.

Compute Performance

PrecisionPeak throughput
FP64 (double precision)34 TFLOPS
FP32 (single precision)67 TFLOPS
FP24134 TFLOPS
FP16 (tensor)1,979 TFLOPS
INT8 (tensor)3,958 TFLOPS
INT4 (tensor)7,916 TFLOPS

Tensor figures use the vendor's peak numbers (with structured sparsity where supported).

System Requirements

Recommended CPU

AMD EPYC 9654 or Intel Xeon Platinum 8490H

Max VRAM per node (8 GPUs)

1,128 GB

System RAM (min / recommended)

512 / 1024 GB

Minimum PSU

1600 W

LLMs on the H200 SXM5

Number of H200 SXM5 GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).

ModelParamsVRAM (8-bit)GPUs needed
GPT-5.6 Sol2400B2682 GB20x H200 SXM5
GPT-5 Flagship2100B2347 GB17x H200 SXM5
GPT-5.6 Luna1400B1565 GB12x H200 SXM5
Kimi K3 (1.2T)1200B1341 GB10x H200 SXM5
Kimi K2.6 (1T)1000B1118 GB8x H200 SXM5
GPT-5.6 Terra800B894 GB7x H200 SXM5
DeepSeek V4 Pro (671B)671B750 GB6x H200 SXM5
Llama 4 Behemoth (500B)500B559 GB4x H200 SXM5
Claude 5 Fable (480B)480B536 GB4x H200 SXM5
GLM 5.2 (400B)400B447 GB4x H200 SXM5
Grok 4.5350B391 GB3x H200 SXM5
Claude 4.8 Opus (300B)300B335 GB3x H200 SXM5
Muse Spark 1.1300B335 GB3x H200 SXM5
Grok 4270B302 GB3x H200 SXM5
Gemini 3.1 Pro250B279 GB2x H200 SXM5
Qwen 3.7 Max (235B)235B263 GB2x H200 SXM5
Mistral Large 3 (200B)200B224 GB2x H200 SXM5
Grok 3 Mini190B212 GB2x H200 SXM5
Claude 5 Sonnet (175B)175B196 GB2x H200 SXM5
Gemini 3.5 Flash150B168 GB2x H200 SXM5
Gemini 2.5 Flash140B156 GB2x H200 SXM5
Llama 4 Maverick (128B)128B143 GB2x H200 SXM5
DeepSeek V4 Flash (120B)120B134 GB1x H200 SXM5
Qwen 3.6 Plus (110B)110B123 GB1x H200 SXM5
Nova Premier (80B)80B89 GB1x H200 SXM5
Qwen 3 Coder-Next (80B)80B89 GB1x H200 SXM5
Claude 4.5 Haiku (70B)70B78 GB1x H200 SXM5
Llama 3.3 Instruct (70B)70B78 GB1x H200 SXM5
Mistral Medium 3.5 (70B)70B78 GB1x H200 SXM5
Yi 1.5 (40B)40B45 GB1x H200 SXM5
Nova Core (34B)34B38 GB1x H200 SXM5
DeepSeek V3.1 (32B)32B36 GB1x H200 SXM5
Gemma 3 (27B)27B30 GB1x H200 SXM5
Mistral Small 4 (24B)24B27 GB1x H200 SXM5
Yi 1.5 (15B)15B17 GB1x H200 SXM5
Phi 4 (14B)14B16 GB1x H200 SXM5
Nova Lite (12B)12B13 GB1x H200 SXM5
Llama 3.2 Instruct (11B)11B12 GB1x H200 SXM5
Gemma 3 (9B)9B10 GB1x H200 SXM5
Yi 1.5 Lite (9B)9B10 GB1x H200 SXM5
Phi 4 Mini (7B)7B8 GB1x H200 SXM5
Phi 3.5 (3.8B)3.8B4 GB1x H200 SXM5

Frequently Asked Questions

How much VRAM does the NVIDIA H200 SXM5 have?

The NVIDIA H200 SXM5 has 141 GB of HBM3e memory with 4800 GB/s of memory bandwidth.

Which LLMs can run on a single H200 SXM5?

At 8-bit quantization, a single H200 SXM5 (141 GB) can serve models up to roughly 120B parameters, such as DeepSeek V4 Flash (120B). Larger models require multiple GPUs or more aggressive quantization.

How much does it cost to rent a NVIDIA H200 SXM5?

On-demand cloud pricing for the H200 SXM5 is around $10.00/hour, i.e. about $7,300/month running 24/7. Actual prices vary by provider, region, and commitment.

What are the power and system requirements of the NVIDIA H200 SXM5?

The H200 SXM5 has a TDP of 700W. A power supply of at least 1600W per GPU is recommended. Recommended host CPUs: AMD EPYC 9654 or Intel Xeon Platinum 8490H.