GPU glossary
The terms you meet when comparing and renting cloud GPUs, in plain English.
Compute
- FLOPS, TFLOPS, PFLOPS
- TeraFLOPS, PetaFLOPS, FLOP/s
- Floating-point operations per second: how much arithmetic a chip can do. 1 TFLOPS is a trillion per second, 1 PFLOPS (PetaFLOPS) is a thousand TFLOPS. PetaPrice shows all compute in PFLOPS.
- Dense vs. sparse FLOPS
- 2:4 sparsity, structured sparsity
- Vendors often quote tensor throughput with 2:4 structured sparsity, which doubles the headline number. Almost no training or inference uses it, so PetaPrice always shows dense figures.
- $ per PFLOP-hour
- $/PFLOP·h, cost per PFLOP
- The price of one GPU-hour divided by the GPU's dense peak PFLOPS at a chosen precision. It shows how much compute a dollar buys, so fast expensive GPUs and slow cheap ones can be compared directly. Lower is better.
- Precision
- number format, data type
- How many bits each number uses. Fewer bits mean more throughput and less memory, at some cost in accuracy. The same GPU has a different peak for each precision.
- FP16 and BF16
- bfloat16, half precision
- 16-bit formats and the standard for training and running LLMs. BF16 keeps FP32's range with less precision, which makes training more stable; it needs Ampere, CDNA 2 or newer.
- FP8
- E4M3, E5M2, Float8
- 8-bit floating point for training and inference, with roughly twice the throughput of FP16. Native on Hopper, Ada, Blackwell and AMD CDNA 3 and newer.
- FP4
- NVFP4, MXFP4
- 4-bit floating point, used for quantized inference. Only Blackwell and AMD CDNA 4 run it natively.
- TF32
- TensorFloat-32
- NVIDIA's 19-bit format that runs FP32 matrix math on tensor cores. PyTorch can use it for FP32 matmuls on Ampere and newer.
- INT8 and TOPS
- 8-bit integer, TOPS
- 8-bit integers for quantized inference. Throughput is counted in TOPS (trillions of operations per second) rather than FLOPS.
- Tensor cores / matrix cores
- matrix cores
- Units that multiply small matrices in one step, where nearly all of a GPU's AI throughput comes from. NVIDIA calls them tensor cores, AMD matrix cores.
- MFU (model FLOPS utilization)
- model FLOPS utilization, GPU utilization
- The share of peak FLOPS a real workload achieves. Large training runs typically reach 30–60%, so datasheet peaks are for comparing GPUs, not for predicting run time.
Memory
- VRAM
- GPU memory
- The GPU's own memory. The model weights, activations and KV cache must fit in it (or be split across GPUs), so VRAM often decides which GPU you can use at all.
- HBM
- HBM2e, HBM3, HBM3e, High Bandwidth Memory
- High Bandwidth Memory, stacked next to the chip on datacenter GPUs (HBM2e, HBM3, HBM3e). Several times faster than the GDDR memory on consumer and workstation cards.
- Memory bandwidth
- GB/s, TB/s
- How fast the GPU reads its memory, in GB/s or TB/s. Generating LLM tokens is usually limited by bandwidth rather than FLOPS, which is why PetaPrice can rank GPUs by bandwidth per dollar.
- KV cache
- key-value cache
- Memory an LLM keeps for every token in its context during inference. Long contexts and many parallel requests need a lot of it, on top of the model weights.
Hardware and interconnect
- SXM vs. PCIe
- SXM5, PCI Express, form factor
- SXM GPUs sit on a baseboard with NVLink between them and get more power, so they run faster. PCIe cards plug into a normal slot, cost less and are usually slower. An H100 SXM and an H100 PCIe are different products.
- NVLink and NVSwitch
- NVSwitch, NVLink-C2C, Infinity Fabric
- NVIDIA's high-speed links between GPUs in one server, far faster than PCIe. They matter when a model is split across GPUs. AMD's equivalent is Infinity Fabric.
- InfiniBand
- IB, RoCE
- The low-latency network used to connect GPU servers in training clusters. Multi-node training needs it (or fast RoCE Ethernet); single-node jobs don't.
- Superchip
- Grace Hopper, Grace Blackwell
- A CPU and GPU on one module with a fast link between them, such as the GH200 and GB200. The GPU can use the CPU's memory as an extension of its own.
- Node
- server, 8-GPU node, HGX
- One server. Datacenter GPUs usually come as 8-GPU nodes; some providers only rent whole nodes. PetaPrice always shows the price per GPU.
- MIG and fractional GPUs
- Multi-Instance GPU, vGPU, fractional GPU
- Multi-Instance GPU splits one A100, H100 or newer into isolated slices. Slices and other fractional GPUs (vGPUs) are cheaper but aren't full GPUs, so PetaPrice leaves them out.
Pricing and renting
- On-demand
- pay-as-you-go
- Pay by the hour (or second) with no commitment and stop whenever you like. The default view on PetaPrice.
- Spot / interruptible
- preemptible, interruptible
- Spare capacity sold at a discount that the provider can take back at any time. Good for checkpointed training and batch jobs.
- Reserved / committed
- committed use, contract pricing
- A lower hourly rate in exchange for a commitment, usually 1 to 3 years or a fixed contract.
- Community / marketplace hosts
- community cloud, peer-to-peer
- Platforms like Vast.ai where independent hosts rent out their machines. Often the cheapest option, with more variation in reliability and security.
- GPU-hour
- GPU·h
- One GPU for one hour. An 8-GPU node for 2 hours is 16 GPU-hours. A GPU running around the clock is about 730 GPU-hours a month.
- Egress
- data transfer out
- Data transferred out of a cloud. Some providers charge for it; PetaPrice prices exclude it, as well as storage and IP addresses, unless the provider bundles them.