GPU Specifications / NVIDIA
NVIDIA H200 SXM5
Hopper datacenter GPU with 141 GB of HBM3e memory, 4800 GB/s of bandwidth and up to 1,979 TFLOPS of FP16 tensor compute.
Key Specifications
Memory
141 GB HBM3e
Memory Bandwidth
4,800 GB/s
TDP
700 W
Architecture
Hopper
Interconnect
NVLink 4.0 · 900 GB/s
Est. On-demand Price
~$10.00/h
Hourly rates are indicative on-demand estimates; actual pricing varies by provider and commitment.
Compute Performance
| Precision | Peak throughput |
|---|---|
| FP64 (double precision) | 34 TFLOPS |
| FP32 (single precision) | 67 TFLOPS |
| FP24 | 134 TFLOPS |
| FP16 (tensor) | 1,979 TFLOPS |
| INT8 (tensor) | 3,958 TFLOPS |
| INT4 (tensor) | 7,916 TFLOPS |
Tensor figures use the vendor's peak numbers (with structured sparsity where supported).
System Requirements
Recommended CPU
AMD EPYC 9654 or Intel Xeon Platinum 8490H
Max VRAM per node (8 GPUs)
1,128 GB
System RAM (min / recommended)
512 / 1024 GB
Minimum PSU
1600 W
LLMs on the H200 SXM5
Number of H200 SXM5 GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).
| Model | Params | VRAM (8-bit) | GPUs needed |
|---|---|---|---|
| GPT-5.6 Sol | 2400B | 2682 GB | 20x H200 SXM5 |
| GPT-5 Flagship | 2100B | 2347 GB | 17x H200 SXM5 |
| GPT-5.6 Luna | 1400B | 1565 GB | 12x H200 SXM5 |
| Kimi K3 (1.2T) | 1200B | 1341 GB | 10x H200 SXM5 |
| Kimi K2.6 (1T) | 1000B | 1118 GB | 8x H200 SXM5 |
| GPT-5.6 Terra | 800B | 894 GB | 7x H200 SXM5 |
| DeepSeek V4 Pro (671B) | 671B | 750 GB | 6x H200 SXM5 |
| Llama 4 Behemoth (500B) | 500B | 559 GB | 4x H200 SXM5 |
| Claude 5 Fable (480B) | 480B | 536 GB | 4x H200 SXM5 |
| GLM 5.2 (400B) | 400B | 447 GB | 4x H200 SXM5 |
| Grok 4.5 | 350B | 391 GB | 3x H200 SXM5 |
| Claude 4.8 Opus (300B) | 300B | 335 GB | 3x H200 SXM5 |
| Muse Spark 1.1 | 300B | 335 GB | 3x H200 SXM5 |
| Grok 4 | 270B | 302 GB | 3x H200 SXM5 |
| Gemini 3.1 Pro | 250B | 279 GB | 2x H200 SXM5 |
| Qwen 3.7 Max (235B) | 235B | 263 GB | 2x H200 SXM5 |
| Mistral Large 3 (200B) | 200B | 224 GB | 2x H200 SXM5 |
| Grok 3 Mini | 190B | 212 GB | 2x H200 SXM5 |
| Claude 5 Sonnet (175B) | 175B | 196 GB | 2x H200 SXM5 |
| Gemini 3.5 Flash | 150B | 168 GB | 2x H200 SXM5 |
| Gemini 2.5 Flash | 140B | 156 GB | 2x H200 SXM5 |
| Llama 4 Maverick (128B) | 128B | 143 GB | 2x H200 SXM5 |
| DeepSeek V4 Flash (120B) | 120B | 134 GB | 1x H200 SXM5 |
| Qwen 3.6 Plus (110B) | 110B | 123 GB | 1x H200 SXM5 |
| Nova Premier (80B) | 80B | 89 GB | 1x H200 SXM5 |
| Qwen 3 Coder-Next (80B) | 80B | 89 GB | 1x H200 SXM5 |
| Claude 4.5 Haiku (70B) | 70B | 78 GB | 1x H200 SXM5 |
| Llama 3.3 Instruct (70B) | 70B | 78 GB | 1x H200 SXM5 |
| Mistral Medium 3.5 (70B) | 70B | 78 GB | 1x H200 SXM5 |
| Yi 1.5 (40B) | 40B | 45 GB | 1x H200 SXM5 |
| Nova Core (34B) | 34B | 38 GB | 1x H200 SXM5 |
| DeepSeek V3.1 (32B) | 32B | 36 GB | 1x H200 SXM5 |
| Gemma 3 (27B) | 27B | 30 GB | 1x H200 SXM5 |
| Mistral Small 4 (24B) | 24B | 27 GB | 1x H200 SXM5 |
| Yi 1.5 (15B) | 15B | 17 GB | 1x H200 SXM5 |
| Phi 4 (14B) | 14B | 16 GB | 1x H200 SXM5 |
| Nova Lite (12B) | 12B | 13 GB | 1x H200 SXM5 |
| Llama 3.2 Instruct (11B) | 11B | 12 GB | 1x H200 SXM5 |
| Gemma 3 (9B) | 9B | 10 GB | 1x H200 SXM5 |
| Yi 1.5 Lite (9B) | 9B | 10 GB | 1x H200 SXM5 |
| Phi 4 Mini (7B) | 7B | 8 GB | 1x H200 SXM5 |
| Phi 3.5 (3.8B) | 3.8B | 4 GB | 1x H200 SXM5 |
Frequently Asked Questions
How much VRAM does the NVIDIA H200 SXM5 have?
The NVIDIA H200 SXM5 has 141 GB of HBM3e memory with 4800 GB/s of memory bandwidth.
Which LLMs can run on a single H200 SXM5?
At 8-bit quantization, a single H200 SXM5 (141 GB) can serve models up to roughly 120B parameters, such as DeepSeek V4 Flash (120B). Larger models require multiple GPUs or more aggressive quantization.
How much does it cost to rent a NVIDIA H200 SXM5?
On-demand cloud pricing for the H200 SXM5 is around $10.00/hour, i.e. about $7,300/month running 24/7. Actual prices vary by provider, region, and commitment.
What are the power and system requirements of the NVIDIA H200 SXM5?
The H200 SXM5 has a TDP of 700W. A power supply of at least 1600W per GPU is recommended. Recommended host CPUs: AMD EPYC 9654 or Intel Xeon Platinum 8490H.
Deploy on a GPU cloud
Rent the NVIDIA H200 SXM5 by the hour instead of buying hardware.