GPU Specifications / NVIDIA
NVIDIA A30
Ampere datacenter GPU with 24 GB of HBM2 memory, 933 GB/s of bandwidth and up to 330 TFLOPS of FP16 tensor compute.
Key Specifications
Memory
24 GB HBM2
Memory Bandwidth
933 GB/s
TDP
165 W
Architecture
Ampere
Interconnect
NVLink 3.0 · 200 GB/s
Est. On-demand Price
~$1.20/h
Hourly rates are indicative on-demand estimates; actual pricing varies by provider and commitment.
Compute Performance
| Precision | Peak throughput |
|---|---|
| FP64 (double precision) | 5.2 TFLOPS |
| FP32 (single precision) | 10.3 TFLOPS |
| FP24 | 20.6 TFLOPS |
| FP16 (tensor) | 330 TFLOPS |
| INT8 (tensor) | 661 TFLOPS |
| INT4 (tensor) | 1,321 TFLOPS |
Tensor figures use the vendor's peak numbers (with structured sparsity where supported).
System Requirements
Recommended CPU
AMD EPYC 7413 or Intel Xeon Gold 6338
Max VRAM per node (8 GPUs)
192 GB
System RAM (min / recommended)
128 / 256 GB
Minimum PSU
600 W
LLMs on the A30
Number of A30 GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).
Frequently Asked Questions
How much VRAM does the NVIDIA A30 have?
The NVIDIA A30 has 24 GB of HBM2 memory with 933 GB/s of memory bandwidth.
Which LLMs can run on a single A30?
At 8-bit quantization, a single A30 (24 GB) can serve models up to roughly 15B parameters, such as Yi 1.5 (15B). Larger models require multiple GPUs or more aggressive quantization.
How much does it cost to rent a NVIDIA A30?
On-demand cloud pricing for the A30 is around $1.20/hour, i.e. about $876/month running 24/7. Actual prices vary by provider, region, and commitment.
What are the power and system requirements of the NVIDIA A30?
The A30 has a TDP of 165W. A power supply of at least 600W per GPU is recommended. Recommended host CPUs: AMD EPYC 7413 or Intel Xeon Gold 6338.
Deploy on a GPU cloud
Rent the NVIDIA A30 by the hour instead of buying hardware.