PetaPrice

GPU glossary

The terms you meet when comparing and renting cloud GPUs, in plain English.

Compute

FLOPS, TFLOPS, PFLOPS
TeraFLOPS, PetaFLOPS, FLOP/s
Floating-point operations per second: how much arithmetic a chip can do. 1 TFLOPS is a trillion per second, 1 PFLOPS (PetaFLOPS) is a thousand TFLOPS. PetaPrice shows all compute in PFLOPS.
Dense vs. sparse FLOPS
2:4 sparsity, structured sparsity
Vendors often quote tensor throughput with 2:4 structured sparsity, which doubles the headline number. Almost no training or inference uses it, so PetaPrice always shows dense figures.
$ per PFLOP-hour
$/PFLOP·h, cost per PFLOP
The price of one GPU-hour divided by the GPU's dense peak PFLOPS at a chosen precision. It shows how much compute a dollar buys, so fast expensive GPUs and slow cheap ones can be compared directly. Lower is better.
Precision
number format, data type
How many bits each number uses. Fewer bits mean more throughput and less memory, at some cost in accuracy. The same GPU has a different peak for each precision.
FP16 and BF16
bfloat16, half precision
16-bit formats and the standard for training and running LLMs. BF16 keeps FP32's range with less precision, which makes training more stable; it needs Ampere, CDNA 2 or newer.
FP8
E4M3, E5M2, Float8
8-bit floating point for training and inference, with roughly twice the throughput of FP16. Native on Hopper, Ada, Blackwell and AMD CDNA 3 and newer.
FP4
NVFP4, MXFP4
4-bit floating point, used for quantized inference. Only Blackwell and AMD CDNA 4 run it natively.
TF32
TensorFloat-32
NVIDIA's 19-bit format that runs FP32 matrix math on tensor cores. PyTorch can use it for FP32 matmuls on Ampere and newer.
INT8 and TOPS
8-bit integer, TOPS
8-bit integers for quantized inference. Throughput is counted in TOPS (trillions of operations per second) rather than FLOPS.
Tensor cores / matrix cores
matrix cores
Units that multiply small matrices in one step, where nearly all of a GPU's AI throughput comes from. NVIDIA calls them tensor cores, AMD matrix cores.
MFU (model FLOPS utilization)
model FLOPS utilization, GPU utilization
The share of peak FLOPS a real workload achieves. Large training runs typically reach 30–60%, so datasheet peaks are for comparing GPUs, not for predicting run time.

Memory

VRAM
GPU memory
The GPU's own memory. The model weights, activations and KV cache must fit in it (or be split across GPUs), so VRAM often decides which GPU you can use at all.
HBM
HBM2e, HBM3, HBM3e, High Bandwidth Memory
High Bandwidth Memory, stacked next to the chip on datacenter GPUs (HBM2e, HBM3, HBM3e). Several times faster than the GDDR memory on consumer and workstation cards.
Memory bandwidth
GB/s, TB/s
How fast the GPU reads its memory, in GB/s or TB/s. Generating LLM tokens is usually limited by bandwidth rather than FLOPS, which is why PetaPrice can rank GPUs by bandwidth per dollar.
KV cache
key-value cache
Memory an LLM keeps for every token in its context during inference. Long contexts and many parallel requests need a lot of it, on top of the model weights.

Hardware and interconnect

SXM vs. PCIe
SXM5, PCI Express, form factor
SXM GPUs sit on a baseboard with NVLink between them and get more power, so they run faster. PCIe cards plug into a normal slot, cost less and are usually slower. An H100 SXM and an H100 PCIe are different products.
InfiniBand
IB, RoCE
The low-latency network used to connect GPU servers in training clusters. Multi-node training needs it (or fast RoCE Ethernet); single-node jobs don't.
Superchip
Grace Hopper, Grace Blackwell
A CPU and GPU on one module with a fast link between them, such as the GH200 and GB200. The GPU can use the CPU's memory as an extension of its own.
Node
server, 8-GPU node, HGX
One server. Datacenter GPUs usually come as 8-GPU nodes; some providers only rent whole nodes. PetaPrice always shows the price per GPU.
TDP
TGP, board power
Thermal design power: roughly the most power the GPU draws, in watts.
MIG and fractional GPUs
Multi-Instance GPU, vGPU, fractional GPU
Multi-Instance GPU splits one A100, H100 or newer into isolated slices. Slices and other fractional GPUs (vGPUs) are cheaper but aren't full GPUs, so PetaPrice leaves them out.

Pricing and renting

On-demand
pay-as-you-go
Pay by the hour (or second) with no commitment and stop whenever you like. The default view on PetaPrice.
Spot / interruptible
preemptible, interruptible
Spare capacity sold at a discount that the provider can take back at any time. Good for checkpointed training and batch jobs.
Reserved / committed
committed use, contract pricing
A lower hourly rate in exchange for a commitment, usually 1 to 3 years or a fixed contract.
Community / marketplace hosts
community cloud, peer-to-peer
Platforms like Vast.ai where independent hosts rent out their machines. Often the cheapest option, with more variation in reliability and security.
GPU-hour
GPU·h
One GPU for one hour. An 8-GPU node for 2 hours is 16 GPU-hours. A GPU running around the clock is about 730 GPU-hours a month.
Egress
data transfer out
Data transferred out of a cloud. Some providers charge for it; PetaPrice prices exclude it, as well as storage and IP addresses, unless the provider bundles them.