Precision
Also known as: number format, data type
How many bits each number uses. Fewer bits mean more throughput and less memory, at some cost in accuracy. The same GPU has a different peak for each precision.
Related terms
- FP16 and BF1616-bit formats and the standard for training and running LLMs. BF16 keeps FP32's range with less precision, which makes training more stable; it needs Ampere, CDNA 2 or newer.
- FP88-bit floating point for training and inference, with roughly twice the throughput of FP16. Native on Hopper, Ada, Blackwell and AMD CDNA 3 and newer.
- FP44-bit floating point, used for quantized inference. Only Blackwell and AMD CDNA 4 run it natively.
- TF32NVIDIA's 19-bit format that runs FP32 matrix math on tensor cores. PyTorch can use it for FP32 matmuls on Ampere and newer.
- INT8 and TOPS8-bit integers for quantized inference. Throughput is counted in TOPS (trillions of operations per second) rather than FLOPS.
Compute
- FLOPS, TFLOPS, PFLOPSFloating-point operations per second: how much arithmetic a chip can do. 1 TFLOPS is a trillion per second, 1 PFLOPS (PetaFLOPS) is a thousand TFLOPS. PetaPrice shows all compute in PFLOPS.
- Dense vs. sparse FLOPSVendors often quote tensor throughput with 2:4 structured sparsity, which doubles the headline number. Almost no training or inference uses it, so PetaPrice always shows dense figures.
- $ per PFLOP-hourThe price of one GPU-hour divided by the GPU's dense peak PFLOPS at a chosen precision. It shows how much compute a dollar buys, so fast expensive GPUs and slow cheap ones can be compared directly. Lower is better.
- Tensor cores / matrix coresUnits that multiply small matrices in one step, where nearly all of a GPU's AI throughput comes from. NVIDIA calls them tensor cores, AMD matrix cores.
- MFU (model FLOPS utilization)The share of peak FLOPS a real workload achieves. Large training runs typically reach 30–60%, so datasheet peaks are for comparing GPUs, not for predicting run time.