…
Cloud rental price, compute per dollar and specs, side by side.
Open in the interactive comparer →| Spec / price | T4 | L4 |
|---|---|---|
| Memory | 16 GB GDDR6 | 24 GB GDDR6better |
| Memory bandwidth | 320 GB/sbetter | 300 GB/s |
| FP16 / BF16 (dense) | 0.065 PF | 0.121 PFbetter |
| FP8 (dense) | — | 0.242 PF |
| FP4 (dense) | — | — |
| INT8 (dense) | 0.13 PF | 0.242 PFbetter |
| TF32 (dense) | — | 0.06 PF |
| FP32 (dense) | 0.0081 PF | 0.0303 PFbetter |
| FP64 (dense) | 0.0003 PF | 0.0005 PFbetter |
| Power | 70 W | 72 W |
| Architecture | Turing (2018) | Ada Lovelace (2023) |
| Lowest on-demand | $0.34 | $0.27better |
| Median on-demand | $0.35better | $0.56 |
| Lowest spot | $0.07better | $0.08 |
| FP16/BF16 $ per PFLOP·h | $5.28 | $2.23better |
| Providers | 3 | 7 |
The L4 has 1.9× more dense FP16/BF16 compute than the T4. The T4 has 1.1× more memory bandwidth than the L4. Real-world speedups depend on the workload: training and prefill scale with compute, while LLM decoding scales mostly with memory bandwidth.
Right now the L4 is cheaper per unit of compute: $2.23 vs $5.28 per PFLOP-hour at the lowest on-demand prices.
Running one GPU around the clock at the lowest on-demand price costs about $250 per month for the T4 and $197 for the L4.
The L4 has more memory: 24 GB vs 16 GB.