Nebius Integrates Inferize Technology Into Token Factory
The acquisition adds engineering work focused on reducing delays when AI models start and scale. The deal’s financial terms were not disclosed.

Nebius is integrating Inferize’s technology and engineering team into its hosted inference platform, Nebius Token Factory, to reduce delays when large AI models start and scale.
Those delays, known as cold starts, can leave graphics processing units (GPUs) idle while a model loads or updates. Inferize’s technology is designed to shorten that wait, allowing computing capacity to adjust more closely to demand and potentially reducing the need for reserve capacity.
Inferize engineers will take part in development and integration work for Token Factory. Nebius has also incorporated technology from Eigen AI for model and core-level tuning, and from Clarifai for inference and computing orchestration.
The Inferize integration focuses on how quickly infrastructure can respond as computing demand changes. The deal’s specific financial terms were not made public.