Turn your neocloud into a premium inference cloud
Premium inference is a new category of AI inference: fast tokens on the largest models, sold at premium prices. It is the fastest-growing demand in the market and the highest-margin tier a neocloud can offer with an ROI in 6 months
SambaNova RDUs are purpose-built to serve AI inference, alongside the GPUs you already run, in the data center you already have.
Talk to Our Neocloud TeamDifferentiate your AI infrastructure with agentic inference
Premium inference is a tier you can price
More than 2×
What the market already pays for high-speed output tokens
The market has determined that speed is worth the price. MiniMax charges more for high-speed output tokens than standard tokens. OpenAI’s Fast mode, Anthropic’s fast mode, and Fireworks’ fast serverless tier all carry a premium over standard processing. Customers are asking for faster ones, not cheaper ones.
~820 TPS
SambaRack SN50 on MiniMax M2.7, interactive configuration
The key to delivering high-speed output tokens lies in handling decode efficiently. Decode is the token-by-token phase every user feels, but GPUs are least suited to run. Adding more GPUs reduces latency a little, but increases power demands, networking complexity, and the margin you want.
13M tokens
Consumed by a single 25-hour OpenAI Codex agent run
SambaNova reconfigurable dataflow unit (RDU) chips are built specifically for the decode phase. This disaggregated inference approach means premium inference stops being an expensive exception in your fleet and becomes a tier on your price list.
Built for agentic AI
Agentic AI is changing what your customers need from you. A chatbot sends one request and stops. An agent plans, calls tools, reads results, generates, checks its work, and loops, sometimes for hours. One recent long-horizon coding agent ran for 25 hours and consumed 13 million tokens on a single task.
Every millisecond of decode latency is multiplied across that loop. Slow tokens do not delay one reply; they delay the entire product your customer is selling. That is why agentic workloads are where premium pricing appeared first, and why they are the workloads worth building capacity.
Powered by RDU chips, SambaStack is purpose-built for agentic inference at scale. The combination of high-speed decode and high throughput is the experience your customers want and the TCO your business needs.
Upgrade your neocloud
Developers want agentic loops to finish in a fraction of the time, and will pay a premium to get there. The challenge for neoclouds is serving tokens fast enough to earn that premium while keeping cost-to-serve low enough to justify it.
Delivering fast tokens is a data-movement problem that SambaNova solved with its Dataflow Architecture. The model graph maps to an efficient path across the processor instead of making repeated trips to off-chip memory. Less memory movement means lower latency and less power for the same work.
The result is what we call the Goldilock’s Zone: fast enough for agents, efficient enough for your margin.
More on RDUs →
Keep your GPUs. Add the decode capacity they’re missing.
Prefill and decode are different jobs. Prefill is compute-bound and highly parallel, so GPUs are a strong fit. Decode is memory-bound and latency-sensitive, which is ideal for RDUs to run. Disaggregated inference splits the pipeline so each stage runs on the architecture built for it.
SambaStack is also available to orchestrate processing between prefill and decode. Models are distributed across your fleet of SambaRack systems and served behind a single OpenAI-compatible API endpoint.
This allows GPUs to stay focused on prefill instead of being over-provisioned for every token-generation path. Utilization goes up across the fleet, not just on the new racks.
Serve the models
your customers
actually want
The most intelligent models are trillions of parameters, and they are the ones commanding premium prices. SambaRack SN50 scales to 256 networked accelerators, supporting models up to 10 trillion parameters and context lengths up to 10 million tokens.
Frontier open models are optimized to run on RDUs, so you can offer the models developers are asking for without a per-model engineering project each time.
More on SambaRack
Neoclouds around
the world
SambaNova powers a growing network of inference clouds and sovereign AI providers delivering fast, efficient, locally-controlled AI inference to their markets. From renewable-powered UK infrastructure to European compliance clouds, Australian onshore AI, and agent-focused neoclouds in the U.S., these providers show what premium inference looks like in production.
Vector Core Compute (VC2)
The world’s first commercially available enterprise inference cloud built on disaggregated inference across CPUs, GPUs, and SambaNova RDUs.
General Compute
A high-performance inference cloud purpose-built for AI agents, including coding and voice agent workloads.
OVHCloud
Fast, energy-efficient inference on leading open-source models for EU developers and enterprises, with a focus on performance, predictable scale, and data sovereignty.
Argyll Data Development
The UK’s first large-scale sovereign AI inference service powered by 100% renewable energy.
SCX AI
An onshore sovereign AI inference cloud for Australian enterprises and government agencies.
Designed for the data
center you already have
Most of the world’s data centers are air-cooled, and moving data for AI workloads is both power-intensive and expensive.
Take the Data Center Walkthrough →
SambaNova’s unique Dataflow Architecture minimizes memory movement on the RDU chip. This energy-saving design allows SambaRack systems to operate within nearly all air-cooled data centers without requiring a liquid-cooled retrofit, a new build, or waiting for a power upgrade.
As a result, SambaRack systems are the only trillion parameter-class inference platform that operates within standard air-cooled power envelopes. It is one of the many reasons sovereign AI inference service providers choose SambaNova.
More on Sovereign AI1. Model the economics
We build the premium-tier revenue and TCO model with you, against your power envelope and your target workloads.
2. Deploy the racks
Turnkey SambaRack systems land in your existing air-cooled facility, integrated by a certified partner.
3. Sell the tier
SambaStack can help to serve your fleet behind one OpenAI-compatible endpoint, so premium inference goes on your price list, not on your roadmap.
RDUs + GPUs co-exist
SambaRack systems are managed seamlessly with SambaStack, the leading hardware and software stack for AI inference. With SambaStack, models are orchestrated across your fleet of SambaRack systems to deliver a standard API end-point on which to run your AI workloads.
SambaStack can also complement your existing GPUs and orchestrate with your existing Kubernetes and inference platforms.
More on SambaStack -->Related resources

SambaNova Launches First Turnkey AI Inference Solution for Data Centers, Deployable in 90 Days
Designed for existing data centers
Most of the world’s data centers today are air-cooled. Data movement running AI workloads can be a power intensive and costly operation.
SambaNova’s unique Dataflow Architecture minimizes memory movement on its RDU chip. This energy-saving design allows SambaRack systems to operate within nearly all air-cooled data centers.
As a result, SambaRack systems are the only solution for power-constrained AI data centers around the world. It is one of the many reasons sovereign AI inference service providers choose SambaNova.
More on sovereign AI -->FAQs
What is premium inference?
Why would a neocloud add RDUs instead of more GPUs?
Because prefill and decode are different problems. GPUs excel at compute-bound prefill; decode is memory-bound and latency-sensitive. Adding GPUs to reduce decode latency raises power draw, networking complexity, and cost-to-serve. RDUs add purpose-built decode capacity, so a neocloud can serve a premium speed tier without over-provisioning the whole fleet.
Do I have to replace my existing GPU fleet?
No. SambaStack is available to orchestrate SambaRack systems alongside your existing GPUs and within your existing inference platform, serving everything behind a single OpenAI-compatible API endpoint. Disaggregated inference routes compute-heavy prefill to GPUs and latency-sensitive decode to RDUs, which improves utilization across the fleet rather than replacing it.
Will SambaRack run in my air-cooled data center?
In nearly all cases, yes. SambaNova’s Dataflow Architecture minimizes memory movement on the RDU chip, which substantially lowers power draw per token. SambaRack systems are designed to operate within standard air-cooled power envelopes, so a power-constrained facility can add trillion-parameter-class inference capacity without a liquid-cooling retrofit or a new build.
How large a model can I serve?
SambaRack SN50 scales to 256 networked accelerators, supporting models up to 10 trillion parameters and context lengths up to 10 million tokens. That covers the frontier open models commanding premium pricing today, with headroom for the next generation.
Does SambaNova compete with its neocloud customers?
SambaNova sells the infrastructure that lets neoclouds serve premium inference themselves. Our business is enabling providers, not competing with them for their customers’ token spend.


