…
Also known as: key-value cache
Memory an LLM keeps for every token in its context during inference. Long contexts and many parallel requests need a lot of it, on top of the model weights.
Ranked by $ per GB of VRAM per hour at each GPU's lowest on-demand price.