Orchestration & Scheduling

Autoscaling (GPU workloads)

Autoscaling automatically adds or removes GPU capacity in response to real-time demand, so infrastructure grows and shrinks with workload rather than sitting at a fixed size.

For GPU infrastructure, autoscaling has to weigh cost against the often significant time it takes to provision a new GPU instance. That tradeoff shapes most orchestration design decisions.

This is one of the core elastic-scaling capabilities of the hosted·ai platform.

Related terms