PetaPrice

Methodology

The numbers

  • $/PFLOP·h = price per GPU-hour ÷ peak PFLOPS at the precision you choose. For example, an H100 SXM at $2.00/h delivers 0.989 dense BF16 PFLOPS, so it costs $2.02 per PFLOP·h. All compute figures on PetaPrice are in PFLOPS.
  • Always dense. Vendors often quote 2:4-sparsity numbers, which double the headline figure. LLM training and inference almost never use structured sparsity, so we halve those numbers. Dense and sparse figures are never mixed.
  • GeForce cards use the FP16/BF16 tensor rate with FP32 accumulate, which is what mixed-precision training actually runs at. NVIDIA markets GeForce with the FP16-accumulate rate, which is twice as high.
  • These are datasheet peaks. Real model FLOPS utilization is usually 30–60%. Use the numbers to compare GPUs with each other, not to predict wall-clock time.
  • Two alternative lenses: $/GB·h of VRAM (what's the cheapest way to fit a model?) and $/TB/s·h of memory bandwidth (LLM decoding speed is usually bandwidth-bound).

The prices

  • Prices come directly from each provider wherever they publish them: public APIs (Vast.ai, RunPod, Verda, Vultr, Cudo, Hyperbolic, Scaleway, OVHcloud), official price lists (AWS, Azure, Oracle Cloud) or their own pricing pages (Lambda, Nebius, Hyperstack, DigitalOcean, Together AI, Crusoe, Hot Aisle, TensorWave, Cirrascale, Seeweb). Shadeform's live multi-cloud API and SkyPilot's public catalogs only fill in providers without a direct source, such as Google Cloud and IBM.
  • The data refreshes every 2 hours. Every price is normalized to USD per GPU per hour. The table also shows the price of the whole instance. Euro prices are converted at the European Central Bank's daily reference rate.
  • On-demand, spot (interruptible) and reserved/committed prices are never mixed. The default view is on-demand only.
  • Hyperscaler SKUs list the cheapest region, with a count of other regions. For marketplaces such as Vast.ai we fold identical listings together and show the cheapest three per configuration.
  • GPUs that providers list but that aren't in our spec sheet yet are shown as "specs pending" with their price and memory, so new hardware appears as soon as it's offered.
  • Prices exclude storage, egress and IP addresses unless the provider bundles them. Vast.ai prices include the host's default disk.

Spec sheet

Peak dense throughput per GPU. INT8 is in TOPS. A dash means the GPU has no native support for that precision. (PFLOPS)

GPUVendorArchitectureVRAMGB/sFP16 / BF16FP8FP4INT8TF32FP32FP64
GB300NVIDIABlackwell Ultra288 GB8,0002.55150.331.250.080.0013
B300 SXMNVIDIABlackwell Ultra288 GB8,0002.254.513.50.331.10.0750.0013
GB200NVIDIABlackwell186 GB8,0002.551051.250.080.045
B200 SXMNVIDIABlackwell180 GB7,7002.254.594.51.10.0750.037
H200 SXMNVIDIAHopper141 GB4,8000.9891.98—1.980.4950.0670.067
H200 NVLNVIDIAHopper141 GB4,8000.8351.67—1.670.4170.060.06
GH200NVIDIAHopper96 GB4,0000.9891.98—1.980.4950.0670.067
H100 SXMNVIDIAHopper80 GB3,3500.9891.98—1.980.4950.0670.067
H100 NVLNVIDIAHopper94 GB3,9000.8351.67—1.670.4170.060.06
H100 PCIeNVIDIAHopper80 GB2,0000.7561.51—1.510.3780.0510.051
H800 SXMNVIDIAHopper80 GB3,3500.9891.98—1.980.4950.0670.001
H20NVIDIAHopper96 GB4,0000.1480.296—0.2960.0740.0440.001
A100 SXM 80GBNVIDIAAmpere80 GB2,0390.312——0.6240.1560.01950.0195
A100 SXM 40GBNVIDIAAmpere40 GB1,5550.312——0.6240.1560.01950.0195
A100 PCIe 80GBNVIDIAAmpere80 GB1,9350.312——0.6240.1560.01950.0195
A100 PCIe 40GBNVIDIAAmpere40 GB1,5550.312——0.6240.1560.01950.0195
A800NVIDIAAmpere80 GB1,9350.312——0.6240.1560.01950.0195
L40SNVIDIAAda Lovelace48 GB8640.3620.733—0.7330.1830.09160.0014
L40NVIDIAAda Lovelace48 GB8640.1810.362—0.3620.09050.09050.0014
L4NVIDIAAda Lovelace24 GB3000.1210.242—0.2420.060.03030.0005
A40NVIDIAAmpere48 GB6960.15——0.2990.07480.03740.0006
A30NVIDIAAmpere24 GB9330.165——0.330.0820.01030.0103
A10NVIDIAAmpere24 GB6000.125——0.250.06250.03120.0005
A10GNVIDIAAmpere24 GB6000.07——0.140.0350.03120.0005
A16NVIDIAAmpere16 GB2000.018——0.0360.0090.0045—
T4NVIDIATuring16 GB3200.065——0.13—0.00810.0003
V100 SXM2 32GBNVIDIAVolta32 GB9000.125————0.01570.0078
V100 SXM2 16GBNVIDIAVolta16 GB9000.125————0.01570.0078
V100 PCIeNVIDIAVolta16 GB9000.112————0.0140.007
V100S PCIeNVIDIAVolta32 GB1,1340.13————0.01640.0082
P100NVIDIAPascal16 GB7320.0187————0.00930.0047
RTX PRO 6000 BlackwellNVIDIABlackwell96 GB1,7920.5041.012.021.010.1260.1260.002
RTX 6000 AdaNVIDIAAda Lovelace48 GB9600.3640.729—0.7290.1820.09110.0014
RTX 5000 AdaNVIDIAAda Lovelace32 GB5760.2610.522—0.5220.1310.06530.001
RTX 4000 AdaNVIDIAAda Lovelace20 GB3600.08180.164—0.1640.04090.02670.0004
RTX A6000NVIDIAAmpere48 GB7680.155——0.310.07740.03870.0006
RTX A5000NVIDIAAmpere24 GB7680.111——0.2220.05560.02780.0004
RTX A4500NVIDIAAmpere20 GB6400.0946——0.1890.04730.02370.0004
RTX A4000NVIDIAAmpere16 GB4480.0767——0.1530.03840.01920.0003
RTX PRO 5000 BlackwellNVIDIABlackwell48 GB1,3440.2680.5351.070.5350.0670.07370.0012
RTX PRO 4000 BlackwellNVIDIABlackwell24 GB6720.1470.2940.5890.2940.0370.04690.0007
RTX 5880 AdaNVIDIAAda Lovelace48 GB9600.2770.555—0.5550.1390.06930.0011
RTX PRO 4500 BlackwellNVIDIABlackwell32 GB8960.2110.4220.8430.4220.0530.05490.0009
RTX 4000 SFF AdaNVIDIAAda Lovelace20 GB2800.07670.153—0.1530.03840.01920.0003
Quadro RTX 6000NVIDIATuring24 GB6720.131——0.261—0.01630.0005
RTX 2000 AdaNVIDIAAda Lovelace16 GB2240.0480.096—0.0960.0240.0120.0002
RTX A2000NVIDIAAmpere12 GB2880.032——0.0640.0160.0080.0001
Quadro RTX 8000NVIDIATuring48 GB6720.131——0.261—0.01630.0005
RTX 5090NVIDIABlackwell32 GB1,7920.2090.4191.680.8380.1050.1050.0016
RTX 5080NVIDIABlackwell16 GB9600.1130.2250.90.450.05630.05630.0009
RTX 5070 TiNVIDIABlackwell16 GB8960.08790.1760.7030.3520.04390.04390.0007
RTX 4090NVIDIAAda Lovelace24 GB1,0080.1650.33—0.6610.08260.08260.0013
RTX 4080 SUPERNVIDIAAda Lovelace16 GB7360.1040.209—0.4180.05220.05220.0008
RTX 4080NVIDIAAda Lovelace16 GB7170.09750.195—0.390.04870.04870.0008
RTX 4070 TiNVIDIAAda Lovelace12 GB5040.08020.16—0.3210.04010.04010.0006
RTX 4070NVIDIAAda Lovelace12 GB5040.05830.117—0.2330.02910.02910.0005
RTX 4060 TiNVIDIAAda Lovelace16 GB2880.04410.0883—0.1770.02210.02210.0003
RTX 3090 TiNVIDIAAmpere24 GB1,0080.08——0.320.040.040.0006
RTX 3090NVIDIAAmpere24 GB9360.071——0.2840.03560.03560.0006
RTX 3080 TiNVIDIAAmpere12 GB9120.0682——0.2730.03410.03410.0005
RTX 3080NVIDIAAmpere10 GB7600.0595——0.2380.02980.02980.0005
RTX 3070NVIDIAAmpere8 GB4480.0406——0.1630.02030.02030.0003
RTX 3060NVIDIAAmpere12 GB3600.0254——0.1020.01270.01270.0002
RTX 2080 TiNVIDIATuring11 GB6160.0538——0.215—0.01340.0004
RTX 4090DNVIDIAAda Lovelace24 GB1,0080.1470.294—0.5880.07350.07350.0011
RTX 5060 TiNVIDIABlackwell16 GB4480.04740.09490.380.190.02370.02370.0004
RTX 4070 SUPERNVIDIAAda Lovelace12 GB5040.0710.142—0.2840.03550.03550.0006
RTX 3070 TiNVIDIAAmpere8 GB6080.0435——0.1740.02170.02170.0003
RTX 3060 TiNVIDIAAmpere8 GB4480.0324——0.130.01620.01620.0003
Titan RTXNVIDIATuring24 GB6720.131——0.261—0.01630.0005
Instinct MI355XAMDCDNA 4288 GB8,0002.525.0310.15.03—0.1570.0786
Instinct MI350XAMDCDNA 4288 GB8,0002.314.619.234.61—0.1440.0721
Instinct MI325XAMDCDNA 3256 GB6,0001.312.61—2.610.6540.1630.163
Instinct MI300XAMDCDNA 3192 GB5,3001.312.61—2.610.6540.1630.163
Instinct MI250XAMDCDNA 2128 GB3,2770.383——0.383—0.09570.0957
Instinct MI250AMDCDNA 2128 GB3,2770.362——0.362—0.09050.0905
Instinct MI210AMDCDNA 264 GB1,6380.181——0.181—0.04530.0453
Instinct MI100AMDCDNA32 GB1,2290.185——0.185—0.04610.0115
Radeon PRO W7900AMDRDNA 348 GB8640.123——0.123—0.06130.0019
Radeon RX 7900 XTXAMDRDNA 324 GB9600.123——0.123—0.06140.0019
Gaudi 3IntelGaudi128 GB3,7001.831.83———0.0287—
Gaudi 2IntelGaudi96 GB2,4500.4320.865———0.011—
Data Center GPU Max 1550IntelXe-HPC128 GB3,2770.839——1.680.4190.05240.0524
TPU v6e (Trillium)Google TPUTPU32 GB1,6400.918——1.84———
TPU v5pGoogle TPUTPU95 GB2,7650.459——0.918———
TPU v5eGoogle TPUTPU16 GB8190.197——0.394———
Trainium2AWS TrainiumTrainium96 GB2,9000.6671.3———0.181—
TrainiumAWS TrainiumTrainium32 GB8200.190.19———0.0475—
Inferentia2AWS TrainiumInferentia32 GB8200.190.19—0.38—0.0475—