Blackwell Ultra
Released 2025NVIDIA's 2025 refresh of Blackwell (B300, GB300): more HBM3e per GPU and faster FP4 for large-scale inference.
Every generation we track, newest first. Each lists its GPUs with today's lowest on-demand price, the precisions its tensor cores run natively and the most memory per GPU.
Hopper vs Blackwell →Most memory and highest power per GPU in each architecture, and its cheapest compute right now.
| Architecture | Vendor | Year | Max memory | Max power | Native precisions | Best FP16 value |
|---|---|---|---|---|---|---|
| Blackwell Ultra | NVIDIA | 2025 | 288 GB HBM3e | 1,400 W | FP4 · FP8 · FP16 / BF16 · TF32 · INT8 | $2.93/PFLOP·h (B300 SXM) |
| Blackwell | NVIDIA | 2024–2025 | 186 GB HBM3e | 1,200 W | FP4 · FP8 · FP16 / BF16 · TF32 · INT8 | $1.42/PFLOP·h (B200 SXM) |
| Hopper | NVIDIA | 2022–2024 | 141 GB HBM3e | 1,000 W | FP8 · FP16 / BF16 · TF32 · INT8 | $1.83/PFLOP·h (H100 SXM) |
| Ada Lovelace | NVIDIA | 2022–2024 | 48 GB GDDR6 | 450 W | FP8 · FP16 / BF16 · TF32 · INT8 | $0.97/PFLOP·h (RTX 4070 SUPER) |
| Ampere | NVIDIA | 2020–2022 | 80 GB HBM2e | 450 W | FP16 / BF16 · TF32 · INT8 | $1.07/PFLOP·h (RTX A4000) |
| Turing | NVIDIA | 2018 | 48 GB GDDR6 | 295 W | FP16 · INT8 | $2.53/PFLOP·h (Quadro RTX 6000) |
| Volta | NVIDIA | 2017–2019 | 32 GB HBM2 | 300 W | FP16 | $1.52/PFLOP·h (V100 SXM2 16GB) |
| Pascal | NVIDIA | 2016 | 16 GB HBM2 | 250 W | FP16 | $68.2/PFLOP·h (P100) |
| CDNA 4 | AMD | 2025 | 288 GB HBM3e | 1,400 W | FP4 · FP8 · FP16 / BF16 · INT8 | $1.03/PFLOP·h (Instinct MI355X) |
| CDNA 3 | AMD | 2023–2024 | 256 GB HBM3e | 1,000 W | FP8 · FP16 / BF16 · TF32 · INT8 | $1.42/PFLOP·h (Instinct MI300X) |
| RDNA 3 | AMD | 2022–2023 | 48 GB GDDR6 | 355 W | FP16 / BF16 · INT8 | — |
| CDNA 2 | AMD | 2021–2022 | 128 GB HBM2e | 560 W | FP16 / BF16 · INT8 | — |
| CDNA | AMD | 2020 | 32 GB HBM2 | 300 W | FP16 / BF16 · INT8 | — |
| Xe-HPC | Intel | 2023 | 128 GB HBM2e | 600 W | FP16 / BF16 · TF32 · INT8 | — |
| Gaudi | Intel | 2022–2024 | 128 GB HBM2e | 900 W | FP8 · FP16 / BF16 | $2.55/PFLOP·h (Gaudi 2) |
| TPU | 2023–2024 | 95 GB HBM2e | — | FP16 / BF16 · INT8 | $2.94/PFLOP·h (TPU v6e (Trillium)) | |
| Inferentia | AWS | 2023 | 32 GB HBM2e | — | FP8 · FP16 / BF16 · INT8 | $3.99/PFLOP·h (Inferentia2) |
| Trainium | AWS | 2022–2024 | 96 GB HBM3 | — | FP8 · FP16 / BF16 | $7.07/PFLOP·h (Trainium) |
NVIDIA's 2025 refresh of Blackwell (B300, GB300): more HBM3e per GPU and faster FP4 for large-scale inference.
NVIDIA's generation after Hopper, from the B200 and GB200 in datacenters to the RTX 50 series and RTX PRO cards. Adds FP4 tensor cores; the datacenter parts join two dies into one GPU.
H100, H200 and GH200. Introduced FP8 tensor cores with the Transformer Engine and is still the most widely rented datacenter GPU.
RTX 40 series, L4, L40S and RTX 6000 Ada. FP8 tensor cores with GDDR6 memory instead of HBM, which makes these cards cheap but bandwidth-limited.
A100, A10, A40 and RTX 30 series. The first generation with BF16 and TF32 tensor cores; the A100 also added MIG partitioning.
T4, RTX 20 series and Quadro RTX. Tensor cores for FP16 and INT8, but no BF16.
V100: the first GPU with tensor cores (FP16 only).
P100: HBM memory but no tensor cores. Mostly useful for cheap FP32 and FP64 work.
AMD Instinct MI350X and MI355X. Adds FP4 and FP6 to AMD's matrix cores, with 288 GB of HBM3e per GPU.
AMD Instinct MI300X and MI325X. FP8 support and more memory per GPU than the Hopper parts they compete with.
AMD's Radeon graphics architecture (RX 7900 XTX, Radeon PRO W7900), rented mostly on community marketplaces.
AMD Instinct MI210 and MI250(X). Strong FP64 for HPC; BF16 but no FP8.
AMD Instinct MI100, AMD's first compute-only datacenter architecture.
Intel Data Center GPU Max (Ponte Vecchio), built for HPC.
Intel's Gaudi 2 and Gaudi 3 AI accelerators (from its Habana Labs acquisition).
Google's own tensor processors, rented only on Google Cloud.
AWS's own inference chips, rented only on AWS.
AWS's own training chips, rented only on AWS.
Ada Lovelace: the RTX 4070 SUPER costs $0.97 per PFLOP-hour of dense FP16/BF16 at $0.07 per GPU-hour. Rankings follow prices, which refresh every 2 hours.
FP8: Blackwell Ultra, Blackwell, Hopper, Ada Lovelace, CDNA 4, CDNA 3, Gaudi, Inferentia, Trainium. FP4: Blackwell Ultra, Blackwell, CDNA 4. Older architectures can still run FP8 or FP4 models, but without the speedup.
Older GPUs are often cheaper per FLOP and per GB of memory, and plenty for serving smaller models, LoRA fine-tuning or experiments. Check that the precision you need runs natively; after that it's a question of price.
Hopper is the GPU architecture (H100, H200). Grace Hopper (GH200) is a superchip: NVIDIA's Arm-based Grace CPU and a Hopper GPU on one module, joined by NVLink-C2C so the GPU can also use the CPU's memory.