GPU Specifications / NVIDIA

NVIDIA A40

Ampere datacenter GPU with 48 GB of GDDR6 memory, 696 GB/s of bandwidth and up to 299.4 TFLOPS of FP16 tensor compute.

Key Specifications

Memory

48 GB GDDR6

Memory Bandwidth

696 GB/s

TDP

300 W

Architecture

Ampere

Interconnect

NVLink 3.0 · 112.5 GB/s

Est. On-demand Price

~$1.80/h

Hourly rates are indicative on-demand estimates; actual pricing varies by provider and commitment.

Compute Performance

PrecisionPeak throughput
FP64 (double precision)0.6 TFLOPS
FP32 (single precision)37.4 TFLOPS
FP2474.8 TFLOPS
FP16 (tensor)299.4 TFLOPS
INT8 (tensor)598.7 TFLOPS
INT4 (tensor)1,197.4 TFLOPS

Tensor figures use the vendor's peak numbers (with structured sparsity where supported).

System Requirements

Recommended CPU

AMD EPYC 7543 or Intel Xeon Gold 6342

Max VRAM per node (8 GPUs)

384 GB

System RAM (min / recommended)

256 / 512 GB

Minimum PSU

850 W

LLMs on the A40

Number of A40 GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).

ModelParamsVRAM (8-bit)GPUs needed
GPT-5.6 Sol2400B2682 GB56x A40
GPT-5 Flagship2100B2347 GB49x A40
GPT-5.6 Luna1400B1565 GB33x A40
Kimi K3 (1.2T)1200B1341 GB28x A40
Kimi K2.6 (1T)1000B1118 GB24x A40
GPT-5.6 Terra800B894 GB19x A40
DeepSeek V4 Pro (671B)671B750 GB16x A40
Llama 4 Behemoth (500B)500B559 GB12x A40
Claude 5 Fable (480B)480B536 GB12x A40
GLM 5.2 (400B)400B447 GB10x A40
Grok 4.5350B391 GB9x A40
Claude 4.8 Opus (300B)300B335 GB7x A40
Muse Spark 1.1300B335 GB7x A40
Grok 4270B302 GB7x A40
Gemini 3.1 Pro250B279 GB6x A40
Qwen 3.7 Max (235B)235B263 GB6x A40
Mistral Large 3 (200B)200B224 GB5x A40
Grok 3 Mini190B212 GB5x A40
Claude 5 Sonnet (175B)175B196 GB5x A40
Gemini 3.5 Flash150B168 GB4x A40
Gemini 2.5 Flash140B156 GB4x A40
Llama 4 Maverick (128B)128B143 GB3x A40
DeepSeek V4 Flash (120B)120B134 GB3x A40
Qwen 3.6 Plus (110B)110B123 GB3x A40
Nova Premier (80B)80B89 GB2x A40
Qwen 3 Coder-Next (80B)80B89 GB2x A40
Claude 4.5 Haiku (70B)70B78 GB2x A40
Llama 3.3 Instruct (70B)70B78 GB2x A40
Mistral Medium 3.5 (70B)70B78 GB2x A40
Yi 1.5 (40B)40B45 GB1x A40
Nova Core (34B)34B38 GB1x A40
DeepSeek V3.1 (32B)32B36 GB1x A40
Gemma 3 (27B)27B30 GB1x A40
Mistral Small 4 (24B)24B27 GB1x A40
Yi 1.5 (15B)15B17 GB1x A40
Phi 4 (14B)14B16 GB1x A40
Nova Lite (12B)12B13 GB1x A40
Llama 3.2 Instruct (11B)11B12 GB1x A40
Gemma 3 (9B)9B10 GB1x A40
Yi 1.5 Lite (9B)9B10 GB1x A40
Phi 4 Mini (7B)7B8 GB1x A40
Phi 3.5 (3.8B)3.8B4 GB1x A40

Frequently Asked Questions

How much VRAM does the NVIDIA A40 have?

The NVIDIA A40 has 48 GB of GDDR6 memory with 696 GB/s of memory bandwidth.

Which LLMs can run on a single A40?

At 8-bit quantization, a single A40 (48 GB) can serve models up to roughly 40B parameters, such as Yi 1.5 (40B). Larger models require multiple GPUs or more aggressive quantization.

How much does it cost to rent a NVIDIA A40?

On-demand cloud pricing for the A40 is around $1.80/hour, i.e. about $1,314/month running 24/7. Actual prices vary by provider, region, and commitment.

What are the power and system requirements of the NVIDIA A40?

The A40 has a TDP of 300W. A power supply of at least 850W per GPU is recommended. Recommended host CPUs: AMD EPYC 7543 or Intel Xeon Gold 6342.