Reconfigurable Dataflow Unit (RDU)

Purpose-built for fast decode and premium AI inference

Designed for agentic workloads

SN50 is SambaNova’s fifth-generation Reconfigurable Dataflow Unit (RDU), purpose-built for the memory-bound phase of inference.

It combines fast token generation, sustained throughput, and efficient data movement to power large intelligent models and AI inference.

Why SN50 is faster at inference

Traditional accelerators often send data back to memory between operations, adding latency and wasted movement. The SN50 RDU chains operations into a continuous dataflow across the processor, reducing repeated memory transfers and relaunch overhead.

The result is faster inference, better throughput, and higher energy efficiency for demanding AI workloads.

traditional-execution
sn50-dataflow

Premium inference for large intelligent models and agents

SN50 is built for interactive AI agents, coding assistants, and multi-step inference workloads that depend on fast token generation, low latency, and sustained throughput.

premium-inference-img-1

Fast
decode

Generates tokens
quickly for premium
inference.

Strong
throughput

Supports concurrent
users and demanding
workloads.

Efficient data
movement

Reduces repeated
movement between
compute and memory.

Agentic
workloads

Designed for coding
and complex inference
applications.


Keep the right data close to compute

SN50 and SN40 use three levels of memory to place data according to how quickly and how often it’s accessed.

  • On-chip SRAM (432 MB) for hot local data, closest to compute, and fastest access
  • HBM2E (64 GB) for active model weights and hot KV cache
  • DDR5 (Up to 512 GB) for prompt caches, larger model pools, and cold KV

This architecture reduces data movement and helps sustain fast token generation across large models and concurrent workloads.

agent-request

From chips to racks

Combining RDU chips into racks powers the largest models with the fast inference. SambaRack systems can be easily integrated into both air-cooled and liquid-cooled data centers. With the SN50 can be used to scale up to 256 RDUs across multiple racks.

More on SambaRack
SN50
SN40
Chips per node
8
16
Compute
1600 TFLOPS at BF16
3200 TFLOPS at FP8
640 TFLOPS
Process node
5 nm TSMC
5 nm TSMC
Architecture
Reconfigurable Dataflow Architecture
Reconfigurable Dataflow Architecture
Memory architecture
Three-tier memory: SRAM, HMB2E, DDR5
Three-tier memory: SRAM, HMB2E, DDR5
On-chip SRAM
432 MB
520 MB
HBM2E
64 GB
64 GB
DDR5
Up to 512 GB
Up to 512 GB
Scale-up
Up to 256 RDUs
Up to 16 RDUs

FAQs

What is the SN50 RDU?

The SN50 RDU (Reconfigurable Dataflow Unit) is SambaNova’s fifth-generation AI inference processor, designed specifically for large-scale, agentic workloads. It uses its unique Dataflow technology and three-tier memory architecture to reduce data movement, enabling faster inference, lower latency, and improved energy efficiency compared to traditional accelerator designs.

How is the RDU different from GPUs?

GPUs are general-purpose accelerators designed to handle a wide range of compute workloads, primarily for training. The SambaNova RDU is purpose-built for inference and uses Dataflow architecture and three-tier memory architecture that maps model execution directly onto the processor, minimizing data movement to memory, which is the most expensive component for AI inference.

What is the difference between SN40 and SN50?

The SN50 is the latest generation of SambaNova’s RDU, offering higher compute performance, increased network bandwidth, and improved scalability compared to the SN40. While the SN40 is well-suited for existing inference deployments and power-constrained environments, the SN50 is designed for large-scale, agentic AI workloads. It enables faster token generation, better system throughput, and more efficient multi-model execution.

What types of AI workloads does the SN50 support?

The SN50 supports a wide range of inference-heavy AI workloads that require low latency, high throughput, and efficient memory usage. These include AI agents, coding assistants, enterprise copilots, conversational AI, retrieval-augmented generation (RAG), and model hosting platforms. It is particularly well-suited to agentic workflows involving multi-step reasoning, tool usage, and frequent model switching.

Can the SN50 run multiple models simultaneously?

Yes, the SN50 is designed to run multiple models simultaneously using its tiered memory architecture. This allows models to remain resident in memory and be switched quickly with minimal latency. The capability is especially important for agentic workloads that rely on multiple models across task steps, improving responsiveness, utilization, and overall inference efficiency.

How does the SN50 scale to large models?

The SN50 scales to large models through a combination of memory architecture and distributed deployment. Multiple racks can be interconnected to form inference clusters, enabling support for larger models, higher concurrency, and predictable performance at scale.

Related Resources

Solving the AI Data Center Power Crisis Without New Construction

Solving the AI Data Center Power Crisis Without New Construction

July 13, 2026
SN50 Runs the Fastest MiniMax Speeds in the World

SN50 Runs the Fastest MiniMax Speeds in the World

July 8, 2026
SambaNova Joins the Genesis Mission Consortium

SambaNova Joins the Genesis Mission Consortium

July 22, 2026