dot point
Platform Updates

‍hosted·ai v2.5: token factories, multi-GPU, and the most advanced GPU scheduler yet

September 8, 2026

We’re excited to announce the availability of hosted·ai v2.5, with a brand new GPU scheduler and new GPUaaS options for neoclouds - now including token factories. Let’s get into it!

What's new in hosted·ai v2.5:

Scheduler v2 - bringing fluidity to GPU infrastructure

Users get better performance and uptime. Providers get better utilization, more scalability, and better manageability for hosting production inference workloads.

The headline feature in this release is a re-engineered GPU scheduler. The scheduler is the heart of the hosted·ai platform: it enables multiple tenants to securely share GPU resources. It turns GPU into an elastic compute resource, solves the problem of idle GPU, and enables neoclouds to make more per GPU hour while delivering lower-cost services for customers.

In the existing scheduler, we provisioned multi-tenant workloads to a pool of GPUs, but each workload was tied to a specific physical card. In hosted·ai v2.5 the scheduler completely decouples AI workloads from physical GPUs and brings a new level of fluidity to multi-tenant GPU infrastructure.  

Now we schedule workloads to any GPU in the pool and continuously load-balance workloads across GPUs, to make optimal use of the resources available.

  • Dynamic scheduling maximizes utilization of GPU resources in the pool, enabling more efficient hosting of multi-tenant inference workloads
  • The scheduler automatically migrates workloads to the most suitable GPU available to optimize performance (if GPUs are becoming overloaded) and to ensure uptime (should a GPU fail, for example)
  • It controls and polices resource access and thresholds: tenants have no direct access to the GPU hardware, improving security and manageability
  • Live migration of workloads also simplifies neocloud operations and SLOs for production inference - in future releases this will be extended to cross-node migration

Multi-GPU support - up to eight GPUs per workload

Host more demanding AI models on each node: multiple tenants can provision up to eight GPUs per workload.

With the new scheduler, up to eight virtualized GPUs can now be assigned to a single workload. This feature works in tandem with hosted·ai’s multi-tenant scheduling capabilities, so that multiple users can map multiple GPUs to workloads simultaneously.

Neocloud providers can now host more demanding AI models oneach node. Multi-GPU support makes use of the NVIDIA Collective Communication Library (NCCL).

GPU emulation - rent out GPUs you didn’t buy

Now customers don’t have to find a new vendor because you didn't have the GPU they wanted.

Customers don’t always need the latest and greatest GPU as a service. They don’t always want to pay a premium for cutting-edge GPU models.

In hosted·ai v2.5, you can present a pool of modern GPUs as any previous generation of card. Now you can serve different customer price points and use cases without buying a large mixed GPU fleet.

  • For example, a pool of NVIDIA GB300s could simultaneously be presented to customers as A100s, H100s or B200s, as well as GB300s
  • Customers pay for the equivalent resources of the smaller card, and you control that pricing
  • Resources are not reserved: they are provisioned on demand by the hosted·ai scheduler

Token Factory - become an inference cloud

Build your own token factory. Host OpenAI-compatible open-weight models, and turn GPUs into token revenue.

Token Factories are available as a new service type in hosted·ai v2.5, alongside bare metal, GPUaaS based on Kubernetes, and VM passthrough based on KVM.

  • Now you can set up public or private tokenfactory services and sell hosted endpoints for a huge range of open-source LLMs, priced per million tokens.
  • Token Factories are fully compatible with the OpenAI API: customers just change a line of code to switch to your service.
  • Token Factories can be built with on-premises GPU or GPU from our capacity trading network (GPU Mesh).‍

You can see the first live deployment of a token factory at packet.ai.

Advanced networking

Private VLAN, VXLAN and IP improvements. More flexible enterprise networking.

Also new in hosted·v2.5: private VLAN capabilities, enabling tunnelling for neocloud customers who need to connect to another data center or cloud environment; VXLAN overlay networks, connecting virtualized GPU instances in a region on a private address range; and dedicated IPs with web service mapping, instead of shared endpoint port redirection, now IPs follow the workload when it moves.

Next steps

This release also includes many smaller features and fixes based on customer feedback.

  • To upgrade from previous versions to hosted·ai v2.5, please contact your account manager or our customer success team.
  • If you're new to hosted·ai, get in touch for a demo and we'll walk/talk you through the platform. Thanks!