Observability & Cost

GPU utilization

GPU utilization is the percentage of a GPU's available compute or memory capacity that's actively being used by workloads at any given time.

Utilization is the core economic metric in GPU infrastructure. Every idle GPU-hour is a sunk cost, whether the hardware is owned or rented. Techniques like fractional GPU, pooling, and overcommit scheduling all exist specifically to push utilization higher.

Related terms

Related resources