Methodology
The numbers
- $/PFLOP·h = price per GPU-hour ÷ peak PFLOPS at the precision you choose. For example, an H100 SXM at $2.00/h delivers 0.989 dense BF16 PFLOPS, so it costs $2.02 per PFLOP·h. All compute figures on PetaPrice are in PFLOPS.
- Always dense. Vendors often quote 2:4-sparsity numbers, which double the headline figure. LLM training and inference almost never use structured sparsity, so we halve those numbers. Dense and sparse figures are never mixed.
- GeForce cards use the FP16/BF16 tensor rate with FP32 accumulate, which is what mixed-precision training actually runs at. NVIDIA markets GeForce with the FP16-accumulate rate, which is twice as high.
- These are datasheet peaks. Real model FLOPS utilization is usually 30–60%. Use the numbers to compare GPUs with each other, not to predict wall-clock time.
- Two alternative lenses: $/GB·h of VRAM (what's the cheapest way to fit a model?) and $/TB/s·h of memory bandwidth (LLM decoding speed is usually bandwidth-bound).
The prices
- Prices come directly from each provider wherever they publish them: public APIs (Vast.ai, RunPod, Verda, Vultr, Cudo, Hyperbolic, Scaleway, OVHcloud), official price lists (AWS, Azure, Oracle Cloud) or their own pricing pages (Lambda, Nebius, Hyperstack, DigitalOcean, Together AI, Crusoe, Hot Aisle, TensorWave, Cirrascale, Seeweb). Shadeform's live multi-cloud API and SkyPilot's public catalogs only fill in providers without a direct source, such as Google Cloud and IBM.
- The data refreshes every 2 hours. Every price is normalized to USD per GPU per hour. The table also shows the price of the whole instance. Euro prices are converted at the European Central Bank's daily reference rate.
- On-demand, spot (interruptible) and reserved/committed prices are never mixed. The default view is on-demand only.
- Hyperscaler SKUs list the cheapest region, with a count of other regions. For marketplaces such as Vast.ai we fold identical listings together and show the cheapest three per configuration.
- GPUs that providers list but that aren't in our spec sheet yet are shown as "specs pending" with their price and memory, so new hardware appears as soon as it's offered.
- Prices exclude storage, egress and IP addresses unless the provider bundles them. Vast.ai prices include the host's default disk.
Spec sheet
Peak dense throughput per GPU. INT8 is in TOPS. A dash means the GPU has no native support for that precision. (PFLOPS)
| GPU | Vendor | Architecture | VRAM | GB/s | FP16 / BF16 | FP8 | FP4 | INT8 | TF32 | FP32 | FP64 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| GB300 | NVIDIA | Blackwell Ultra | 288 GB | 8,000 | 2.5 | 5 | 15 | 0.33 | 1.25 | 0.08 | 0.0013 |
| B300 SXM | NVIDIA | Blackwell Ultra | 288 GB | 8,000 | 2.25 | 4.5 | 13.5 | 0.33 | 1.1 | 0.075 | 0.0013 |
| GB200 | NVIDIA | Blackwell | 186 GB | 8,000 | 2.5 | 5 | 10 | 5 | 1.25 | 0.08 | 0.045 |
| B200 SXM | NVIDIA | Blackwell | 180 GB | 7,700 | 2.25 | 4.5 | 9 | 4.5 | 1.1 | 0.075 | 0.037 |
| H200 SXM | NVIDIA | Hopper | 141 GB | 4,800 | 0.989 | 1.98 | — | 1.98 | 0.495 | 0.067 | 0.067 |
| H200 NVL | NVIDIA | Hopper | 141 GB | 4,800 | 0.835 | 1.67 | — | 1.67 | 0.417 | 0.06 | 0.06 |
| GH200 | NVIDIA | Hopper | 96 GB | 4,000 | 0.989 | 1.98 | — | 1.98 | 0.495 | 0.067 | 0.067 |
| H100 SXM | NVIDIA | Hopper | 80 GB | 3,350 | 0.989 | 1.98 | — | 1.98 | 0.495 | 0.067 | 0.067 |
| H100 NVL | NVIDIA | Hopper | 94 GB | 3,900 | 0.835 | 1.67 | — | 1.67 | 0.417 | 0.06 | 0.06 |
| H100 PCIe | NVIDIA | Hopper | 80 GB | 2,000 | 0.756 | 1.51 | — | 1.51 | 0.378 | 0.051 | 0.051 |
| H800 SXM | NVIDIA | Hopper | 80 GB | 3,350 | 0.989 | 1.98 | — | 1.98 | 0.495 | 0.067 | 0.001 |
| H20 | NVIDIA | Hopper | 96 GB | 4,000 | 0.148 | 0.296 | — | 0.296 | 0.074 | 0.044 | 0.001 |
| A100 SXM 80GB | NVIDIA | Ampere | 80 GB | 2,039 | 0.312 | — | — | 0.624 | 0.156 | 0.0195 | 0.0195 |
| A100 SXM 40GB | NVIDIA | Ampere | 40 GB | 1,555 | 0.312 | — | — | 0.624 | 0.156 | 0.0195 | 0.0195 |
| A100 PCIe 80GB | NVIDIA | Ampere | 80 GB | 1,935 | 0.312 | — | — | 0.624 | 0.156 | 0.0195 | 0.0195 |
| A100 PCIe 40GB | NVIDIA | Ampere | 40 GB | 1,555 | 0.312 | — | — | 0.624 | 0.156 | 0.0195 | 0.0195 |
| A800 | NVIDIA | Ampere | 80 GB | 1,935 | 0.312 | — | — | 0.624 | 0.156 | 0.0195 | 0.0195 |
| L40S | NVIDIA | Ada Lovelace | 48 GB | 864 | 0.362 | 0.733 | — | 0.733 | 0.183 | 0.0916 | 0.0014 |
| L40 | NVIDIA | Ada Lovelace | 48 GB | 864 | 0.181 | 0.362 | — | 0.362 | 0.0905 | 0.0905 | 0.0014 |
| L4 | NVIDIA | Ada Lovelace | 24 GB | 300 | 0.121 | 0.242 | — | 0.242 | 0.06 | 0.0303 | 0.0005 |
| A40 | NVIDIA | Ampere | 48 GB | 696 | 0.15 | — | — | 0.299 | 0.0748 | 0.0374 | 0.0006 |
| A30 | NVIDIA | Ampere | 24 GB | 933 | 0.165 | — | — | 0.33 | 0.082 | 0.0103 | 0.0103 |
| A10 | NVIDIA | Ampere | 24 GB | 600 | 0.125 | — | — | 0.25 | 0.0625 | 0.0312 | 0.0005 |
| A10G | NVIDIA | Ampere | 24 GB | 600 | 0.07 | — | — | 0.14 | 0.035 | 0.0312 | 0.0005 |
| A16 | NVIDIA | Ampere | 16 GB | 200 | 0.018 | — | — | 0.036 | 0.009 | 0.0045 | — |
| T4 | NVIDIA | Turing | 16 GB | 320 | 0.065 | — | — | 0.13 | — | 0.0081 | 0.0003 |
| V100 SXM2 32GB | NVIDIA | Volta | 32 GB | 900 | 0.125 | — | — | — | — | 0.0157 | 0.0078 |
| V100 SXM2 16GB | NVIDIA | Volta | 16 GB | 900 | 0.125 | — | — | — | — | 0.0157 | 0.0078 |
| V100 PCIe | NVIDIA | Volta | 16 GB | 900 | 0.112 | — | — | — | — | 0.014 | 0.007 |
| V100S PCIe | NVIDIA | Volta | 32 GB | 1,134 | 0.13 | — | — | — | — | 0.0164 | 0.0082 |
| P100 | NVIDIA | Pascal | 16 GB | 732 | 0.0187 | — | — | — | — | 0.0093 | 0.0047 |
| RTX PRO 6000 Blackwell | NVIDIA | Blackwell | 96 GB | 1,792 | 0.504 | 1.01 | 2.02 | 1.01 | 0.126 | 0.126 | 0.002 |
| RTX 6000 Ada | NVIDIA | Ada Lovelace | 48 GB | 960 | 0.364 | 0.729 | — | 0.729 | 0.182 | 0.0911 | 0.0014 |
| RTX 5000 Ada | NVIDIA | Ada Lovelace | 32 GB | 576 | 0.261 | 0.522 | — | 0.522 | 0.131 | 0.0653 | 0.001 |
| RTX 4000 Ada | NVIDIA | Ada Lovelace | 20 GB | 360 | 0.0818 | 0.164 | — | 0.164 | 0.0409 | 0.0267 | 0.0004 |
| RTX A6000 | NVIDIA | Ampere | 48 GB | 768 | 0.155 | — | — | 0.31 | 0.0774 | 0.0387 | 0.0006 |
| RTX A5000 | NVIDIA | Ampere | 24 GB | 768 | 0.111 | — | — | 0.222 | 0.0556 | 0.0278 | 0.0004 |
| RTX A4500 | NVIDIA | Ampere | 20 GB | 640 | 0.0946 | — | — | 0.189 | 0.0473 | 0.0237 | 0.0004 |
| RTX A4000 | NVIDIA | Ampere | 16 GB | 448 | 0.0767 | — | — | 0.153 | 0.0384 | 0.0192 | 0.0003 |
| RTX PRO 5000 Blackwell | NVIDIA | Blackwell | 48 GB | 1,344 | 0.268 | 0.535 | 1.07 | 0.535 | 0.067 | 0.0737 | 0.0012 |
| RTX PRO 4000 Blackwell | NVIDIA | Blackwell | 24 GB | 672 | 0.147 | 0.294 | 0.589 | 0.294 | 0.037 | 0.0469 | 0.0007 |
| RTX 5880 Ada | NVIDIA | Ada Lovelace | 48 GB | 960 | 0.277 | 0.555 | — | 0.555 | 0.139 | 0.0693 | 0.0011 |
| RTX PRO 4500 Blackwell | NVIDIA | Blackwell | 32 GB | 896 | 0.211 | 0.422 | 0.843 | 0.422 | 0.053 | 0.0549 | 0.0009 |
| RTX 4000 SFF Ada | NVIDIA | Ada Lovelace | 20 GB | 280 | 0.0767 | 0.153 | — | 0.153 | 0.0384 | 0.0192 | 0.0003 |
| Quadro RTX 6000 | NVIDIA | Turing | 24 GB | 672 | 0.131 | — | — | 0.261 | — | 0.0163 | 0.0005 |
| RTX 2000 Ada | NVIDIA | Ada Lovelace | 16 GB | 224 | 0.048 | 0.096 | — | 0.096 | 0.024 | 0.012 | 0.0002 |
| RTX A2000 | NVIDIA | Ampere | 12 GB | 288 | 0.032 | — | — | 0.064 | 0.016 | 0.008 | 0.0001 |
| Quadro RTX 8000 | NVIDIA | Turing | 48 GB | 672 | 0.131 | — | — | 0.261 | — | 0.0163 | 0.0005 |
| RTX 5090 | NVIDIA | Blackwell | 32 GB | 1,792 | 0.209 | 0.419 | 1.68 | 0.838 | 0.105 | 0.105 | 0.0016 |
| RTX 5080 | NVIDIA | Blackwell | 16 GB | 960 | 0.113 | 0.225 | 0.9 | 0.45 | 0.0563 | 0.0563 | 0.0009 |
| RTX 5070 Ti | NVIDIA | Blackwell | 16 GB | 896 | 0.0879 | 0.176 | 0.703 | 0.352 | 0.0439 | 0.0439 | 0.0007 |
| RTX 4090 | NVIDIA | Ada Lovelace | 24 GB | 1,008 | 0.165 | 0.33 | — | 0.661 | 0.0826 | 0.0826 | 0.0013 |
| RTX 4080 SUPER | NVIDIA | Ada Lovelace | 16 GB | 736 | 0.104 | 0.209 | — | 0.418 | 0.0522 | 0.0522 | 0.0008 |
| RTX 4080 | NVIDIA | Ada Lovelace | 16 GB | 717 | 0.0975 | 0.195 | — | 0.39 | 0.0487 | 0.0487 | 0.0008 |
| RTX 4070 Ti | NVIDIA | Ada Lovelace | 12 GB | 504 | 0.0802 | 0.16 | — | 0.321 | 0.0401 | 0.0401 | 0.0006 |
| RTX 4070 | NVIDIA | Ada Lovelace | 12 GB | 504 | 0.0583 | 0.117 | — | 0.233 | 0.0291 | 0.0291 | 0.0005 |
| RTX 4060 Ti | NVIDIA | Ada Lovelace | 16 GB | 288 | 0.0441 | 0.0883 | — | 0.177 | 0.0221 | 0.0221 | 0.0003 |
| RTX 3090 Ti | NVIDIA | Ampere | 24 GB | 1,008 | 0.08 | — | — | 0.32 | 0.04 | 0.04 | 0.0006 |
| RTX 3090 | NVIDIA | Ampere | 24 GB | 936 | 0.071 | — | — | 0.284 | 0.0356 | 0.0356 | 0.0006 |
| RTX 3080 Ti | NVIDIA | Ampere | 12 GB | 912 | 0.0682 | — | — | 0.273 | 0.0341 | 0.0341 | 0.0005 |
| RTX 3080 | NVIDIA | Ampere | 10 GB | 760 | 0.0595 | — | — | 0.238 | 0.0298 | 0.0298 | 0.0005 |
| RTX 3070 | NVIDIA | Ampere | 8 GB | 448 | 0.0406 | — | — | 0.163 | 0.0203 | 0.0203 | 0.0003 |
| RTX 3060 | NVIDIA | Ampere | 12 GB | 360 | 0.0254 | — | — | 0.102 | 0.0127 | 0.0127 | 0.0002 |
| RTX 2080 Ti | NVIDIA | Turing | 11 GB | 616 | 0.0538 | — | — | 0.215 | — | 0.0134 | 0.0004 |
| RTX 4090D | NVIDIA | Ada Lovelace | 24 GB | 1,008 | 0.147 | 0.294 | — | 0.588 | 0.0735 | 0.0735 | 0.0011 |
| RTX 5060 Ti | NVIDIA | Blackwell | 16 GB | 448 | 0.0474 | 0.0949 | 0.38 | 0.19 | 0.0237 | 0.0237 | 0.0004 |
| RTX 4070 SUPER | NVIDIA | Ada Lovelace | 12 GB | 504 | 0.071 | 0.142 | — | 0.284 | 0.0355 | 0.0355 | 0.0006 |
| RTX 3070 Ti | NVIDIA | Ampere | 8 GB | 608 | 0.0435 | — | — | 0.174 | 0.0217 | 0.0217 | 0.0003 |
| RTX 3060 Ti | NVIDIA | Ampere | 8 GB | 448 | 0.0324 | — | — | 0.13 | 0.0162 | 0.0162 | 0.0003 |
| Titan RTX | NVIDIA | Turing | 24 GB | 672 | 0.131 | — | — | 0.261 | — | 0.0163 | 0.0005 |
| Instinct MI355X | AMD | CDNA 4 | 288 GB | 8,000 | 2.52 | 5.03 | 10.1 | 5.03 | — | 0.157 | 0.0786 |
| Instinct MI350X | AMD | CDNA 4 | 288 GB | 8,000 | 2.31 | 4.61 | 9.23 | 4.61 | — | 0.144 | 0.0721 |
| Instinct MI325X | AMD | CDNA 3 | 256 GB | 6,000 | 1.31 | 2.61 | — | 2.61 | 0.654 | 0.163 | 0.163 |
| Instinct MI300X | AMD | CDNA 3 | 192 GB | 5,300 | 1.31 | 2.61 | — | 2.61 | 0.654 | 0.163 | 0.163 |
| Instinct MI250X | AMD | CDNA 2 | 128 GB | 3,277 | 0.383 | — | — | 0.383 | — | 0.0957 | 0.0957 |
| Instinct MI250 | AMD | CDNA 2 | 128 GB | 3,277 | 0.362 | — | — | 0.362 | — | 0.0905 | 0.0905 |
| Instinct MI210 | AMD | CDNA 2 | 64 GB | 1,638 | 0.181 | — | — | 0.181 | — | 0.0453 | 0.0453 |
| Instinct MI100 | AMD | CDNA | 32 GB | 1,229 | 0.185 | — | — | 0.185 | — | 0.0461 | 0.0115 |
| Radeon PRO W7900 | AMD | RDNA 3 | 48 GB | 864 | 0.123 | — | — | 0.123 | — | 0.0613 | 0.0019 |
| Radeon RX 7900 XTX | AMD | RDNA 3 | 24 GB | 960 | 0.123 | — | — | 0.123 | — | 0.0614 | 0.0019 |
| Gaudi 3 | Intel | Gaudi | 128 GB | 3,700 | 1.83 | 1.83 | — | — | — | 0.0287 | — |
| Gaudi 2 | Intel | Gaudi | 96 GB | 2,450 | 0.432 | 0.865 | — | — | — | 0.011 | — |
| Data Center GPU Max 1550 | Intel | Xe-HPC | 128 GB | 3,277 | 0.839 | — | — | 1.68 | 0.419 | 0.0524 | 0.0524 |
| TPU v6e (Trillium) | Google TPU | TPU | 32 GB | 1,640 | 0.918 | — | — | 1.84 | — | — | — |
| TPU v5p | Google TPU | TPU | 95 GB | 2,765 | 0.459 | — | — | 0.918 | — | — | — |
| TPU v5e | Google TPU | TPU | 16 GB | 819 | 0.197 | — | — | 0.394 | — | — | — |
| Trainium2 | AWS Trainium | Trainium | 96 GB | 2,900 | 0.667 | 1.3 | — | — | — | 0.181 | — |
| Trainium | AWS Trainium | Trainium | 32 GB | 820 | 0.19 | 0.19 | — | — | — | 0.0475 | — |
| Inferentia2 | AWS Trainium | Inferentia | 32 GB | 820 | 0.19 | 0.19 | — | 0.38 | — | 0.0475 | — |