Nebius adds Alibaba's 27B Qwen model to its hosted inference menu
Nebius said the open-weight model was live on Token Factory on September 28th; Alibaba's Qwen team highlighted the availability two days later.
By Ryan Merket · Published
Primary source: Qwen / Alibaba
Why it matters
The post is about distribution, not a new model: Nebius adds a managed route to open Qwen weights, while leaving price, performance and customer uptake unanswered.

Qwen3.8-27B became available through Nebius Token Factory's hosted inference endpoint after Nebius said the model was live on the service on September 28th. Alibaba's Qwen team highlighted the access route on September 30th. The model's weights had been available since August.
Qwen supplies the model; Nebius sells the infrastructure and managed inference used to run it. For teams testing an open-weight model without operating their own serving stack, a shared hosted endpoint removes an infrastructure step. Nebius's endpoint page labels the shared API as suitable for tests, not production, and notes that availability may change. The listing does not establish that this option is cheaper or faster than running the model elsewhere.
Nebius founder and CEO Arkady Volozh, who previously led Yandex, described Nebius's broader ambition in 2024 as building a full-stack AI infrastructure provider. Adding access to another open model fits that stated direction: Nebius is competing to serve the systems around AI development, as well as the compute underneath them. The announcement does not attribute this particular addition to Volozh or include a comment from him.
A model already available to deploy
Qwen3.8-27B was already distributed as open weights before the Nebius post. Its model card records availability on Hugging Face and ModelScope on August 14th, 2026. The card describes it as a 27-billion-parameter dense vision-language model, with a native context window of 262,144 tokens that can be extended to 1 million. It supports image and video understanding as well as text, and offers configurable thinking behavior.
The model card also documents deployment through tools including vLLM and SGLang. Those routes give developers the option to serve the weights themselves; a hosted endpoint trades some infrastructure control for a managed service. Qwen's September 30th post frames the model for agent workflows and deep research, but offers no benchmark results or workload-specific evidence to show how well it handles those tasks on Nebius.
Qwen3.8-27B had already launched; the newly reported development is its availability on Nebius Token Factory. Nebius launched the platform in 2025 to host open and custom models. At launch, Nebius said it offered OpenAI-compatible APIs and supported models including Qwen. The September posts identify this specific Qwen model as live on the service.
A wider hosted catalog gives Nebius more options for teams choosing among open models. Qwen gets another managed endpoint for developers who do not want to handle model-serving operations themselves. Nebius's endpoint listing displays an approximate price of $0.45 per million input tokens and $3 per million output tokens, with displayed throughput of 218 tokens per second. The page gives no customer counts or usage data.