…
Cloud rental price, compute per dollar and specs, side by side.
Open in the interactive comparer →| Spec / price | L40S | H100 PCIe |
|---|---|---|
| Memory | 48 GB GDDR6 | 80 GB HBM2ebetter |
| Memory bandwidth | 864 GB/s | 2,000 GB/sbetter |
| FP16 / BF16 (dense) | 0.362 PF | 0.756 PFbetter |
| FP8 (dense) | 0.733 PF | 1.51 PFbetter |
| FP4 (dense) | — | — |
| INT8 (dense) | 0.733 PF | 1.51 PFbetter |
| TF32 (dense) | 0.183 PF | 0.378 PFbetter |
| FP32 (dense) | 0.0916 PFbetter | 0.051 PF |
| FP64 (dense) | 0.0014 PF | 0.051 PFbetter |
| Power | 350 W | 350 W |
| Architecture | Ada Lovelace (2023) | Hopper (2022) |
| Lowest on-demand | $0.60better | $1.98 |
| Median on-demand | $1.50better | $2.75 |
| Lowest spot | $0.38better | $2.00 |
| FP16/BF16 $ per PFLOP·h | $1.66better | $2.62 |
| Providers | 17 | 6 |
The H100 PCIe has 2.1× more dense FP16/BF16 compute than the L40S. The H100 PCIe has 2.3× more memory bandwidth than the L40S. Real-world speedups depend on the workload: training and prefill scale with compute, while LLM decoding scales mostly with memory bandwidth.
Right now the L40S is cheaper per unit of compute: $1.66 vs $2.62 per PFLOP-hour at the lowest on-demand prices.
Running one GPU around the clock at the lowest on-demand price costs about $440 per month for the L40S and $1,445 for the H100 PCIe.
The H100 PCIe has more memory: 80 GB vs 48 GB.