We’re excited to announce the availability of hosted·ai v2.5, with a brand new GPU scheduler and new GPUaaS options for neoclouds - now including token factories. Let’s get into it!
Users get better performance and uptime. Providers get better utilization, more scalability, and better manageability for hosting production inference workloads.
The headline feature in this release is a re-engineered GPU scheduler. The scheduler is the heart of the hosted·ai platform: it enables multiple tenants to securely share GPU resources. It turns GPU into an elastic compute resource, solves the problem of idle GPU, and enables neoclouds to make more per GPU hour while delivering lower-cost services for customers.
In the existing scheduler, we provisioned multi-tenant workloads to a pool of GPUs, but each workload was tied to a specific physical card. In hosted·ai v2.5 the scheduler completely decouples AI workloads from physical GPUs and brings a new level of fluidity to multi-tenant GPU infrastructure.
Now we schedule workloads to any GPU in the pool and continuously load-balance workloads across GPUs, to make optimal use of the resources available.
Host more demanding AI models on each node: multiple tenants can provision up to eight GPUs per workload.
With the new scheduler, up to eight virtualized GPUs can now be assigned to a single workload. This feature works in tandem with hosted·ai’s multi-tenant scheduling capabilities, so that multiple users can map multiple GPUs to workloads simultaneously.
Neocloud providers can now host more demanding AI models oneach node. Multi-GPU support makes use of the NVIDIA Collective Communication Library (NCCL).
Now customers don’t have to find a new vendor because you didn't have the GPU they wanted.
Customers don’t always need the latest and greatest GPU as a service. They don’t always want to pay a premium for cutting-edge GPU models.
In hosted·ai v2.5, you can present a pool of modern GPUs as any previous generation of card. Now you can serve different customer price points and use cases without buying a large mixed GPU fleet.
Build your own token factory. Host OpenAI-compatible open-weight models, and turn GPUs into token revenue.
Token Factories are available as a new service type in hosted·ai v2.5, alongside bare metal, GPUaaS based on Kubernetes, and VM passthrough based on KVM.
You can see the first live deployment of a token factory at packet.ai.
Private VLAN, VXLAN and IP improvements. More flexible enterprise networking.
Also new in hosted·v2.5: private VLAN capabilities, enabling tunnelling for neocloud customers who need to connect to another data center or cloud environment; VXLAN overlay networks, connecting virtualized GPU instances in a region on a private address range; and dedicated IPs with web service mapping, instead of shared endpoint port redirection, now IPs follow the workload when it moves.
This release also includes many smaller features and fixes based on customer feedback.
We use cookies for analytics and advertising. Privacy policy