…
Cloud rental price, compute per dollar and specs, side by side.
Open in the interactive comparer →| Spec / price | A10 | L4 |
|---|---|---|
| Memory | 24 GB GDDR6 | 24 GB GDDR6 |
| Memory bandwidth | 600 GB/sbetter | 300 GB/s |
| FP16 / BF16 (dense) | 0.125 PFbetter | 0.121 PF |
| FP8 (dense) | — | 0.242 PF |
| FP4 (dense) | — | — |
| INT8 (dense) | 0.25 PFbetter | 0.242 PF |
| TF32 (dense) | 0.0625 PFbetter | 0.06 PF |
| FP32 (dense) | 0.0312 PFbetter | 0.0303 PF |
| FP64 (dense) | 0.0005 PF | 0.0005 PF |
| Power | 150 W | 72 W |
| Architecture | Ampere (2021) | Ada Lovelace (2023) |
| Lowest on-demand | $1.29 | $0.27better |
| Median on-demand | $2.00 | $0.49better |
| Lowest spot | $0.59 | $0.08better |
| FP16/BF16 $ per PFLOP·h | $10.3 | $2.23better |
| Providers | 3 | 5 |
The A10 has 1.0× more dense FP16/BF16 compute than the L4. The A10 has 2.0× more memory bandwidth than the L4. Real-world speedups depend on the workload: training and prefill scale with compute, while LLM decoding scales mostly with memory bandwidth.
Right now the L4 is cheaper per unit of compute: $2.23 vs $10.3 per PFLOP-hour at the lowest on-demand prices.
Running one GPU around the clock at the lowest on-demand price costs about $942 per month for the A10 and $197 for the L4.
Both have 24 GB of VRAM; the A10 has more bandwidth.