Turn your neocloud into a premium inference cloud

Premium inference is a new category of AI inference: fast tokens on the largest models, sold at premium prices. It is the fastest-growing demand in the market and the highest-margin tier a neocloud can offer with an ROI in 6 months

See the SN50 benchmarks ⟶

SambaNova RDUs are purpose-built to serve AI inference, alongside the GPUs you already run, in the data center you already have.

Talk to Our Neocloud Team
Inference Providers

Differentiate your AI infrastructure with agentic inference

Premium inference is a tier you can price

More than 2×

What the market already pays for high-speed output tokens

Logo - MiniMax

The market has determined that speed is worth the price. MiniMax charges more for high-speed output tokens than standard tokens. OpenAI’s Fast mode, Anthropic’s fast mode, and Fireworks’ fast serverless tier all carry a premium over standard processing. Customers are asking for faster ones, not cheaper ones.

~820 TPS

SambaRack SN50 on MiniMax M2.7, interactive configuration

Logo - semianalysis

The key to delivering high-speed output tokens lies in handling decode efficiently. Decode is the token-by-token phase every user feels, but GPUs are least suited to run. Adding more GPUs reduces latency a little, but increases power demands, networking complexity, and the margin you want.

13M tokens

Consumed by a single 25-hour OpenAI Codex agent run

Logo - OpenAI

SambaNova reconfigurable dataflow unit (RDU) chips are built specifically for the decode phase. This disaggregated inference approach means premium inference stops being an expensive exception in your fleet and becomes a tier on your price list.


Built for agentic AI

Agentic AI is changing what your customers need from you. A chatbot sends one request and stops. An agent plans, calls tools, reads results, generates, checks its work, and loops, sometimes for hours. One recent long-horizon coding agent ran for 25 hours and consumed 13 million tokens on a single task.

Every millisecond of decode latency is multiplied across that loop. Slow tokens do not delay one reply; they delay the entire product your customer is selling. That is why agentic workloads are where premium pricing appeared first, and why they are the workloads worth building capacity.

Powered by RDU chips, SambaStack is purpose-built for agentic inference at scale. The combination of high-speed decode and high throughput is the experience your customers want and the TCO your business needs.

Upgrade your neocloud

SambaNova - SN50 CHIP - 600w

Developers want agentic loops to finish in a fraction of the time, and will pay a premium to get there. The challenge for neoclouds is serving tokens fast enough to earn that premium while keeping cost-to-serve low enough to justify it.

Delivering fast tokens is a data-movement problem that SambaNova solved with its Dataflow Architecture. The model graph maps to an efficient path across the processor instead of making repeated trips to off-chip memory. Less memory movement means lower latency and less power for the same work.

The result is what we call the Goldilock’s Zone: fast enough for agents, efficient enough for your margin.

More on RDUs →

 

SambaNova - SN50 CHIP - 600w

Keep your GPUs. Add the decode capacity they’re missing.

Prefill and decode are different jobs. Prefill is compute-bound and highly parallel, so GPUs are a strong fit. Decode is memory-bound and latency-sensitive, which is ideal for RDUs to run. Disaggregated inference splits the pipeline so each stage runs on the architecture built for it.

SambaStack is also available to orchestrate processing between prefill and decode. Models are distributed across your fleet of SambaRack systems and served behind a single OpenAI-compatible API endpoint.

This allows GPUs to stay focused on prefill instead of being over-provisioned for every token-generation path. Utilization goes up across the fleet, not just on the new racks.

Serve the models
your customers
actually want

The most intelligent models are trillions of parameters, and they are the ones commanding premium prices. SambaRack SN50 scales to 256 networked accelerators, supporting models up to 10 trillion parameters and context lengths up to 10 million tokens.

Frontier open models are optimized to run on RDUs, so you can offer the models developers are asking for without a per-model engineering project each time.

More on SambaRack
SambaNova - Serve the models your customers actually want - Image

Neoclouds around
the world

SambaNova powers a growing network of inference clouds and sovereign AI providers delivering fast, efficient, locally-controlled AI inference to their markets. From renewable-powered UK infrastructure to European compliance clouds, Australian onshore AI, and agent-focused neoclouds in the U.S., these providers show what premium inference looks like in production.

Designed for the data
center you already have

Most of the world’s data centers are air-cooled, and moving data for AI workloads is both power-intensive and expensive.

Take the Data Center Walkthrough →

SambaNova’s unique Dataflow Architecture minimizes memory movement on the RDU chip. This energy-saving design allows SambaRack systems to operate within nearly all air-cooled data centers without requiring a liquid-cooled retrofit, a new build, or waiting for a power upgrade.

As a result, SambaRack systems are the only trillion parameter-class inference platform that operates within standard air-cooled power envelopes. It is one of the many reasons sovereign AI inference service providers choose SambaNova.

More on Sovereign AI

3 simple steps for fast
deployment

SambaRack ships as a turnkey inference system, deployable in existing data centers. SambaStack is available to help deliver the API endpoint your customers connect to.

1. Model the economics

We build the premium-tier revenue and TCO model with you, against your power envelope and your target workloads.

2. Deploy the racks

Turnkey SambaRack systems land in your existing air-cooled facility, integrated by a certified partner.

3. Sell the tier

SambaStack can help to serve your fleet behind one OpenAI-compatible endpoint, so premium inference goes on your price list, not on your roadmap.


2026 05 18 - Chart Inference AP1 - v2.1

RDUs + GPUs co-exist

SambaRack systems are managed seamlessly with SambaStack, the leading hardware and software stack for AI inference. With SambaStack, models are orchestrated across your fleet of SambaRack systems to deliver a standard API end-point on which to run your AI workloads.

SambaStack can also complement your existing GPUs and orchestrate with your existing Kubernetes and inference platforms.

More on SambaStack -->

Related resources

Inference Speed or Throughput? With RDUs, You Don't Have to Choose

Inference Speed or Throughput? With RDUs, You Don't Have to Choose

January 15, 2026
SambaNova Launches First Turnkey AI Inference Solution for Data Centers, Deployable in 90 Days

SambaNova Launches First Turnkey AI Inference Solution for Data Centers, Deployable in 90 Days

July 7, 2025
SambaNova Launches its AI Platform in AWS Marketplace

SambaNova Launches its AI Platform in AWS Marketplace

May 29, 2025

Designed for existing data centers

Most of the world’s data centers today are air-cooled. Data movement running AI workloads can be a power intensive and costly operation.

SambaNova’s unique Dataflow Architecture minimizes memory movement on its RDU chip. This energy-saving design allows SambaRack systems to operate within nearly all air-cooled data centers.

As a result, SambaRack systems are the only solution for power-constrained AI data centers around the world. It is one of the many reasons sovereign AI inference service providers choose SambaNova.

More on sovereign AI -->

FAQs

What is premium inference?

Premium inference is fast, responsive serving of large, intelligent models, offered as a distinct paid tier rather than a single commodity endpoint. It matters most for agentic workloads, where the model generates in a loop and every millisecond of decode latency is multiplied across the entire task. Providers including OpenAI, Anthropic, MiniMax, and Fireworks already price it above standard serving.

Why would a neocloud add RDUs instead of more GPUs?

Because prefill and decode are different problems. GPUs excel at compute-bound prefill; decode is memory-bound and latency-sensitive. Adding GPUs to reduce decode latency raises power draw, networking complexity, and cost-to-serve. RDUs add purpose-built decode capacity, so a neocloud can serve a premium speed tier without over-provisioning the whole fleet.

Do I have to replace my existing GPU fleet?

No. SambaStack is available to orchestrate SambaRack systems alongside your existing GPUs and within your existing inference platform, serving everything behind a single OpenAI-compatible API endpoint. Disaggregated inference routes compute-heavy prefill to GPUs and latency-sensitive decode to RDUs, which improves utilization across the fleet rather than replacing it.

Will SambaRack run in my air-cooled data center?

In nearly all cases, yes. SambaNova’s Dataflow Architecture minimizes memory movement on the RDU chip, which substantially lowers power draw per token. SambaRack systems are designed to operate within standard air-cooled power envelopes, so a power-constrained facility can add trillion-parameter-class inference capacity without a liquid-cooling retrofit or a new build.

How large a model can I serve?

 SambaRack SN50 scales to 256 networked accelerators, supporting models up to 10 trillion parameters and context lengths up to 10 million tokens. That covers the frontier open models commanding premium pricing today, with headroom for the next generation.

Does SambaNova compete with its neocloud customers?

SambaNova sells the infrastructure that lets neoclouds serve premium inference themselves. Our business is enabling providers, not competing with them for their customers’ token spend.