GPU Specifications / NVIDIA
NVIDIA A10
Ampere datacenter GPU with 24 GB of GDDR6 memory, 600 GB/s of bandwidth and up to 250 TFLOPS of FP16 tensor compute.
Key Specifications
Memory
24 GB GDDR6
Memory Bandwidth
600 GB/s
TDP
150 W
Architecture
Ampere
Interconnect
PCIe 4.0 · 64 GB/s
Est. On-demand Price
~$1.00/h
Hourly rates are indicative on-demand estimates; actual pricing varies by provider and commitment.
Compute Performance
| Precision | Peak throughput |
|---|---|
| FP64 (double precision) | 0.5 TFLOPS |
| FP32 (single precision) | 31.2 TFLOPS |
| FP24 | 62.4 TFLOPS |
| FP16 (tensor) | 250 TFLOPS |
| INT8 (tensor) | 500 TFLOPS |
| INT4 (tensor) | 1,000 TFLOPS |
Tensor figures use the vendor's peak numbers (with structured sparsity where supported).
System Requirements
Recommended CPU
AMD EPYC 7413 or Intel Xeon Gold 6338
Max VRAM per node (8 GPUs)
192 GB
System RAM (min / recommended)
128 / 256 GB
Minimum PSU
600 W
LLMs on the A10
Number of A10 GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).
Frequently Asked Questions
How much VRAM does the NVIDIA A10 have?
The NVIDIA A10 has 24 GB of GDDR6 memory with 600 GB/s of memory bandwidth.
Which LLMs can run on a single A10?
At 8-bit quantization, a single A10 (24 GB) can serve models up to roughly 15B parameters, such as Yi 1.5 (15B). Larger models require multiple GPUs or more aggressive quantization.
How much does it cost to rent a NVIDIA A10?
On-demand cloud pricing for the A10 is around $1.00/hour, i.e. about $730/month running 24/7. Actual prices vary by provider, region, and commitment.
What are the power and system requirements of the NVIDIA A10?
The A10 has a TDP of 150W. A power supply of at least 600W per GPU is recommended. Recommended host CPUs: AMD EPYC 7413 or Intel Xeon Gold 6338.
Deploy on a GPU cloud
Rent the NVIDIA A10 by the hour instead of buying hardware.