Frequently asked questions
How PetaPrice works, where the numbers come from and how to pick a GPU.
Prices and numbers
What is $ per PFLOP-hour?
It is the rental price of one GPU for one hour divided by its peak compute in PetaFLOPS at the precision you choose. It shows how much compute each dollar buys, so a cheap slow GPU and an expensive fast one can be compared directly. Lower is better.
Which GPU gives the most compute per dollar right now?
At dense FP16/BF16, the Instinct MI355X currently leads at $1.03 per PFLOP-hour ($2.59 per GPU-hour at Vultr). Rankings change with prices; the table above updates every 2 hours.
How much does an H100 cost per hour?
On-demand H100 SXM prices currently start at $1.82 per GPU-hour, with a median of $3.56 across 21 providers. Spot capacity starts at $1.25.
Why dense FLOPS instead of the numbers on NVIDIA's spec sheets?
Headline tensor numbers often assume 2:4 structured sparsity, which doubles them but is rarely used in practice. PetaPrice always uses dense throughput so GPUs and vendors are compared fairly.
Where does the price data come from?
Directly from the providers wherever they publish prices (public APIs, official price lists and pricing pages), with Shadeform's multi-cloud API and SkyPilot's catalogs filling gaps: 871 offers from 32 providers, refreshed about every 2 hours. Prices are indicative list prices; confirm with the provider before booking.
What's the difference between on-demand, spot and reserved?
On-demand is billed by the hour with no commitment. Spot (interruptible) capacity is cheaper but can be reclaimed at any time. Reserved prices require a commitment, typically 1 to 3 years. PetaPrice never mixes them; the default view is on-demand.
Choosing a GPU
Which GPU should I rent for LLM inference?
First make sure the model fits: weights plus KV cache must fit in VRAM (a 70B model in FP16 needs about 140 GB for weights alone). Then compare by memory bandwidth, because generating tokens is usually bandwidth-bound; switch “Optimize for” to Bandwidth on the home page. For training and long-prompt prefill, compare FLOPS instead.
Which GPU is best for training?
Compare $ per PFLOP-hour at the precision you train in (usually BF16, or FP8 on Hopper and newer). For multi-GPU training, prefer SXM GPUs with NVLink; across several servers you also need InfiniBand. Spot capacity is much cheaper if your job checkpoints regularly.
About the data
Why does a provider's website show a different price?
Prices change, differ by region, and some providers bill CPU, RAM or storage separately. PetaPrice refreshes about every 2 hours and shows the cheapest region; the provider's own website is authoritative when you book. If a price is consistently wrong, please report it.
Why don't I see MIG slices or fractional GPUs?
A slice of a GPU isn't comparable with a whole one, so MIG slices, vGPUs and time-shared GPUs are left out everywhere. Every price on PetaPrice is for a full GPU.
What does “community” mean?
Community hosts rent out machines through peer-to-peer marketplaces such as Vast.ai or RunPod Community Cloud. They are often the cheapest, but reliability and security vary by host. Use the Host filter to show only managed clouds.
What does the EU filter show?
Providers headquartered in an EU member state. It filters by where the company is based, not where its GPUs run, so check the provider's regions if data residency matters.
Is there an API?
Yes. /api/offers returns every normalized offer plus the GPU spec sheet as JSON, with CORS enabled. /llms.txt summarizes the site for AI assistants.
I found a wrong price. How do I report it?
Send us the GPU, the provider and a link to its pricing page. The Corrections page explains how, and lists what we've fixed so far.