PetaPrice

GPU architectures

Every generation we track, newest first. Each lists its GPUs with today's lowest on-demand price, the precisions its tensor cores run natively and the most memory per GPU.

Hopper vs Blackwell →

Architecture, product, superchip

Architecture
The design a family of chips shares: tensor cores, supported precisions, memory system. Hopper, Blackwell and CDNA 3 are architectures.
Product
One GPU built on an architecture, in one memory size and form factor. H100 SXM, H100 PCIe and H200 are three products on Hopper, each with its own price.
Superchip
A CPU and one or two GPUs on one module (GH200, GB200). PetaPrice shows the price per GPU; the CPU comes with it.

Side by side

Most memory and highest power per GPU in each architecture, and its cheapest compute right now.

ArchitectureVendorYearMax memoryMax powerNative precisionsBest FP16 value
Blackwell UltraNVIDIA2025288 GB HBM3e1,400 WFP4 · FP8 · FP16 / BF16 · TF32 · INT8$2.93/PFLOP·h (B300 SXM)
BlackwellNVIDIA2024–2025186 GB HBM3e1,200 WFP4 · FP8 · FP16 / BF16 · TF32 · INT8$1.42/PFLOP·h (B200 SXM)
HopperNVIDIA2022–2024141 GB HBM3e1,000 WFP8 · FP16 / BF16 · TF32 · INT8$1.83/PFLOP·h (H100 SXM)
Ada LovelaceNVIDIA2022–202448 GB GDDR6450 WFP8 · FP16 / BF16 · TF32 · INT8$0.97/PFLOP·h (RTX 4070 SUPER)
AmpereNVIDIA2020–202280 GB HBM2e450 WFP16 / BF16 · TF32 · INT8$1.07/PFLOP·h (RTX A4000)
TuringNVIDIA201848 GB GDDR6295 WFP16 · INT8$2.53/PFLOP·h (Quadro RTX 6000)
VoltaNVIDIA2017–201932 GB HBM2300 WFP16$1.52/PFLOP·h (V100 SXM2 16GB)
PascalNVIDIA201616 GB HBM2250 WFP16$68.2/PFLOP·h (P100)
CDNA 4AMD2025288 GB HBM3e1,400 WFP4 · FP8 · FP16 / BF16 · INT8$1.03/PFLOP·h (Instinct MI355X)
CDNA 3AMD2023–2024256 GB HBM3e1,000 WFP8 · FP16 / BF16 · TF32 · INT8$1.42/PFLOP·h (Instinct MI300X)
RDNA 3AMD2022–202348 GB GDDR6355 WFP16 / BF16 · INT8—
CDNA 2AMD2021–2022128 GB HBM2e560 WFP16 / BF16 · INT8—
CDNAAMD202032 GB HBM2300 WFP16 / BF16 · INT8—
Xe-HPCIntel2023128 GB HBM2e600 WFP16 / BF16 · TF32 · INT8—
GaudiIntel2022–2024128 GB HBM2e900 WFP8 · FP16 / BF16$2.55/PFLOP·h (Gaudi 2)
TPUGoogle2023–202495 GB HBM2e—FP16 / BF16 · INT8$2.94/PFLOP·h (TPU v6e (Trillium))
InferentiaAWS202332 GB HBM2e—FP8 · FP16 / BF16 · INT8$3.99/PFLOP·h (Inferentia2)
TrainiumAWS2022–202496 GB HBM3—FP8 · FP16 / BF16$7.07/PFLOP·h (Trainium)

NVIDIA

Blackwell Ultra

Released 2025

NVIDIA's 2025 refresh of Blackwell (B300, GB300): more HBM3e per GPU and faster FP4 for large-scale inference.

Up to 288 GB per GPUNative precisions: FP4 · FP8 · FP16 / BF16 · TF32 · INT82 of 2 rentable now

Blackwell

Released 2024–2025

NVIDIA's generation after Hopper, from the B200 and GB200 in datacenters to the RTX 50 series and RTX PRO cards. Adds FP4 tensor cores; the datacenter parts join two dies into one GPU.

Up to 186 GB per GPUNative precisions: FP4 · FP8 · FP16 / BF16 · TF32 · INT810 of 10 rentable now

Ada Lovelace

Released 2022–2024

RTX 40 series, L4, L40S and RTX 6000 Ada. FP8 tensor cores with GDDR6 memory instead of HBM, which makes these cards cheap but bandwidth-limited.

Up to 48 GB per GPUNative precisions: FP8 · FP16 / BF16 · TF32 · INT815 of 17 rentable now

Ampere

Released 2020–2022

A100, A10, A40 and RTX 30 series. The first generation with BF16 and TF32 tensor cores; the A100 also added MIG partitioning.

Up to 80 GB per GPUNative precisions: FP16 / BF16 · TF32 · INT819 of 23 rentable now

Pascal

Released 2016

P100: HBM memory but no tensor cores. Mostly useful for cheap FP32 and FP64 work.

Up to 16 GB per GPUNative precisions: FP161 of 1 rentable now

AMD

CDNA 3

Released 2023–2024

AMD Instinct MI300X and MI325X. FP8 support and more memory per GPU than the Hopper parts they compete with.

Up to 256 GB per GPUNative precisions: FP8 · FP16 / BF16 · TF32 · INT82 of 2 rentable now

CDNA

Released 2020

AMD Instinct MI100, AMD's first compute-only datacenter architecture.

Up to 32 GB per GPUNative precisions: FP16 / BF16 · INT80 of 1 rentable now

Intel

Xe-HPC

Released 2023

Intel Data Center GPU Max (Ponte Vecchio), built for HPC.

Up to 128 GB per GPUNative precisions: FP16 / BF16 · TF32 · INT80 of 1 rentable now

Gaudi

Released 2022–2024

Intel's Gaudi 2 and Gaudi 3 AI accelerators (from its Habana Labs acquisition).

Up to 128 GB per GPUNative precisions: FP8 · FP16 / BF161 of 2 rentable now

Google

AWS

Inferentia

Released 2023

AWS's own inference chips, rented only on AWS.

Up to 32 GB per GPUNative precisions: FP8 · FP16 / BF16 · INT81 of 1 rentable now

Architecture questions

Which architecture gives the most compute per dollar right now?

Ada Lovelace: the RTX 4070 SUPER costs $0.97 per PFLOP-hour of dense FP16/BF16 at $0.07 per GPU-hour. Rankings follow prices, which refresh every 2 hours.

Which architectures support FP8 and FP4?

FP8: Blackwell Ultra, Blackwell, Hopper, Ada Lovelace, CDNA 4, CDNA 3, Gaudi, Inferentia, Trainium. FP4: Blackwell Ultra, Blackwell, CDNA 4. Older architectures can still run FP8 or FP4 models, but without the speedup.

Why rent an older architecture?

Older GPUs are often cheaper per FLOP and per GB of memory, and plenty for serving smaller models, LoRA fine-tuning or experiments. Check that the precision you need runs natively; after that it's a question of price.

What's the difference between Hopper and Grace Hopper?

Hopper is the GPU architecture (H100, H200). Grace Hopper (GH200) is a superchip: NVIDIA's Arm-based Grace CPU and a Hopper GPU on one module, joined by NVLink-C2C so the GPU can also use the CPU's memory.