GPU Specifications / NVIDIA

NVIDIA L40S

Ada Lovelace datacenter GPU with 48 GB of GDDR6 memory, 864 GB/s of bandwidth and up to 733 TFLOPS of FP16 tensor compute.

Key Specifications

Memory

48 GB GDDR6

Memory Bandwidth

864 GB/s

TDP

350 W

Architecture

Ada Lovelace

Interconnect

PCIe 4.0 · 64 GB/s

Est. On-demand Price

~$3.50/h

Hourly rates are indicative on-demand estimates; actual pricing varies by provider and commitment.

Compute Performance

PrecisionPeak throughput
FP64 (double precision)1.4 TFLOPS
FP32 (single precision)91.6 TFLOPS
FP24183.2 TFLOPS
FP16 (tensor)733 TFLOPS
INT8 (tensor)1,466 TFLOPS
INT4 (tensor)2,932 TFLOPS

Tensor figures use the vendor's peak numbers (with structured sparsity where supported).

System Requirements

Recommended CPU

AMD EPYC 7443 or Intel Xeon Gold 6348

Max VRAM per node (8 GPUs)

384 GB

System RAM (min / recommended)

128 / 256 GB

Minimum PSU

800 W

LLMs on the L40S

Number of L40S GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).

ModelParamsVRAM (8-bit)GPUs needed
GPT-5.6 Sol2400B2682 GB56x L40S
GPT-5 Flagship2100B2347 GB49x L40S
GPT-5.6 Luna1400B1565 GB33x L40S
Kimi K3 (1.2T)1200B1341 GB28x L40S
Kimi K2.6 (1T)1000B1118 GB24x L40S
GPT-5.6 Terra800B894 GB19x L40S
DeepSeek V4 Pro (671B)671B750 GB16x L40S
Llama 4 Behemoth (500B)500B559 GB12x L40S
Claude 5 Fable (480B)480B536 GB12x L40S
GLM 5.2 (400B)400B447 GB10x L40S
Grok 4.5350B391 GB9x L40S
Claude 4.8 Opus (300B)300B335 GB7x L40S
Muse Spark 1.1300B335 GB7x L40S
Grok 4270B302 GB7x L40S
Gemini 3.1 Pro250B279 GB6x L40S
Qwen 3.7 Max (235B)235B263 GB6x L40S
Mistral Large 3 (200B)200B224 GB5x L40S
Grok 3 Mini190B212 GB5x L40S
Claude 5 Sonnet (175B)175B196 GB5x L40S
Gemini 3.5 Flash150B168 GB4x L40S
Gemini 2.5 Flash140B156 GB4x L40S
Llama 4 Maverick (128B)128B143 GB3x L40S
DeepSeek V4 Flash (120B)120B134 GB3x L40S
Qwen 3.6 Plus (110B)110B123 GB3x L40S
Nova Premier (80B)80B89 GB2x L40S
Qwen 3 Coder-Next (80B)80B89 GB2x L40S
Claude 4.5 Haiku (70B)70B78 GB2x L40S
Llama 3.3 Instruct (70B)70B78 GB2x L40S
Mistral Medium 3.5 (70B)70B78 GB2x L40S
Yi 1.5 (40B)40B45 GB1x L40S
Nova Core (34B)34B38 GB1x L40S
DeepSeek V3.1 (32B)32B36 GB1x L40S
Gemma 3 (27B)27B30 GB1x L40S
Mistral Small 4 (24B)24B27 GB1x L40S
Yi 1.5 (15B)15B17 GB1x L40S
Phi 4 (14B)14B16 GB1x L40S
Nova Lite (12B)12B13 GB1x L40S
Llama 3.2 Instruct (11B)11B12 GB1x L40S
Gemma 3 (9B)9B10 GB1x L40S
Yi 1.5 Lite (9B)9B10 GB1x L40S
Phi 4 Mini (7B)7B8 GB1x L40S
Phi 3.5 (3.8B)3.8B4 GB1x L40S

Frequently Asked Questions

How much VRAM does the NVIDIA L40S have?

The NVIDIA L40S has 48 GB of GDDR6 memory with 864 GB/s of memory bandwidth.

Which LLMs can run on a single L40S?

At 8-bit quantization, a single L40S (48 GB) can serve models up to roughly 40B parameters, such as Yi 1.5 (40B). Larger models require multiple GPUs or more aggressive quantization.

How much does it cost to rent a NVIDIA L40S?

On-demand cloud pricing for the L40S is around $3.50/hour, i.e. about $2,555/month running 24/7. Actual prices vary by provider, region, and commitment.

What are the power and system requirements of the NVIDIA L40S?

The L40S has a TDP of 350W. A power supply of at least 800W per GPU is recommended. Recommended host CPUs: AMD EPYC 7443 or Intel Xeon Gold 6348.