Infra & Ops

SLA (GPU infrastructure)

An SLA (service-level agreement) is a formal commitment on GPU infrastructure performance or availability, such as uptime, provisioning time, or guaranteed compute access, with defined consequences if it isn't met.

GPU-specific SLAs often need to address things generic cloud SLAs don't, like guaranteed VRAM availability or maximum preemption frequency for reserved capacity.

Related terms