Nebius Token FactoryNebius
|
||||||
Related Products
|
||||||
About
OpenAI- and Anthropic-compatible inference API from an EU company. The flagship model runs on dedicated GPUs in EIA data centres with zero data retention: prompts and completions are processed in memory only, not stored, not logged, not used for training. Routed open models from third-party providers are available with the same key and clearly labelled. One DPA and one invoice from an EU company. Features: streaming, tool calling, structured output, public DPA and sub-processor list, per-token pricing. Measured on the live system in August 2026: 176 tokens per second per stream, first token in 0.3 seconds.
|
About
Nebius Token Factory is a scalable AI inference platform designed to run open-source and custom AI models in production without manual infrastructure management. It offers enterprise-ready inference endpoints with predictable performance, autoscaling throughput, and sub-second latency — even at very high request volumes. It delivers 99.9% uptime availability and supports unlimited or tailored traffic profiles based on workload needs, simplifying the transition from experimentation to global deployment. Nebius Token Factory supports a broad set of open source models such as Llama, Qwen, DeepSeek, GPT-OSS, Flux, and many others, and lets teams host and fine-tune models through an API or dashboard. Users can upload LoRA adapters or full fine-tuned variants directly, with the same enterprise performance guarantees applied to custom models.
|
|||||
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
|||||
Audience
Companies that need LLM inference under EU data-protection requirements; developers of AI agents and RAG applications
|
Audience
Engineering and data science teams that need a production-grade inference system to deploy, scale, and manage open-source or custom AI models reliably in enterprise environments
|
|||||
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Supported
|
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Supported
|
|||||
API
Offers API
Not Supported
|
API
Offers API
Supported
|
|||||
Screenshots and VideosNo images available
|
Screenshots and Videos |
|||||
Pricing
$0.04 per 1M input tokens
Free Version
Supported
Free Trial
Not Supported
|
Pricing
$0.02
Free Version
Supported
Free Trial
Supported
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Not Supported
In Person
Not Supported
|
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Not Supported
In Person
Not Supported
|
|||||
Company InformationHeabsy
Founded: 2014
Slovakia
heabsy.com
|
Company InformationNebius
Founded: 2022
Netherlands
nebius.com/services/token-factory/enterprise-grade-inference
|
|||||
Alternatives |
Alternatives |
|||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
Categories |
Categories |
|||||
Integrations
BGE
Not Supported
DeepSeek V3.1
Not Supported
Devstral Small 2
Not Supported
FLUX.1
Not Supported
GLM-4.5
Not Supported
Gemma 3
Not Supported
Kimi
Not Supported
Kimi K2 Thinking
Not Supported
Llama
Not Supported
Llama 3.3
Not Supported
|
Integrations
BGE
Supported
DeepSeek V3.1
Supported
Devstral Small 2
Supported
FLUX.1
Supported
GLM-4.5
Supported
Gemma 3
Supported
Kimi
Supported
Kimi K2 Thinking
Supported
Llama
Supported
Llama 3.3
Supported
|
|||||
|
|
|