Reconfigurable Dataflow Unit (RDU)
Purpose-built for fast decode and premium AI inference
Designed for agentic workloads
SN50 is SambaNova’s fifth-generation Reconfigurable Dataflow Unit (RDU), purpose-built for the memory-bound phase of inference.
It combines fast token generation, sustained throughput, and efficient data movement to power large intelligent models and AI inference.
Why SN50 is faster at inference
Traditional accelerators often send data back to memory between operations, adding latency and wasted movement. The SN50 RDU chains operations into a continuous dataflow across the processor, reducing repeated memory transfers and relaunch overhead.
The result is faster inference, better throughput, and higher energy efficiency for demanding AI workloads.
Premium inference for large intelligent models and agents
SN50 is built for interactive AI agents, coding assistants, and multi-step inference workloads that depend on fast token generation, low latency, and sustained throughput.
Fast
decode
Generates tokens
quickly for premium
inference.
Strong
throughput
Supports concurrent
users and demanding
workloads.
Efficient data
movement
Reduces repeated
movement between
compute and memory.
Agentic
workloads
Designed for coding
and complex inference
applications.
Keep the right data close to compute
SN50 and SN40 use three levels of memory to place data according to how quickly and how often it’s accessed.
- On-chip SRAM (432 MB) for hot local data, closest to compute, and fastest access
- HBM2E (64 GB) for active model weights and hot KV cache
- DDR5 (Up to 512 GB) for prompt caches, larger model pools, and cold KV
This architecture reduces data movement and helps sustain fast token generation across large models and concurrent workloads.
From chips to racks
Combining RDU chips into racks powers the largest models with the fast inference. SambaRack systems can be easily integrated into both air-cooled and liquid-cooled data centers. With the SN50 can be used to scale up to 256 RDUs across multiple racks.
More on SambaRackSN50 and SN40 RDU Specifications
3200 TFLOPS at FP8
FAQs
The SN50 RDU (Reconfigurable Dataflow Unit) is SambaNova’s fifth-generation AI inference processor, designed specifically for large-scale, agentic workloads. It uses its unique Dataflow technology and three-tier memory architecture to reduce data movement, enabling faster inference, lower latency, and improved energy efficiency compared to traditional accelerator designs.
GPUs are general-purpose accelerators designed to handle a wide range of compute workloads, primarily for training. The SambaNova RDU is purpose-built for inference and uses Dataflow architecture and three-tier memory architecture that maps model execution directly onto the processor, minimizing data movement to memory, which is the most expensive component for AI inference.
The SN50 is the latest generation of SambaNova’s RDU, offering higher compute performance, increased network bandwidth, and improved scalability compared to the SN40. While the SN40 is well-suited for existing inference deployments and power-constrained environments, the SN50 is designed for large-scale, agentic AI workloads. It enables faster token generation, better system throughput, and more efficient multi-model execution.
The SN50 supports a wide range of inference-heavy AI workloads that require low latency, high throughput, and efficient memory usage. These include AI agents, coding assistants, enterprise copilots, conversational AI, retrieval-augmented generation (RAG), and model hosting platforms. It is particularly well-suited to agentic workflows involving multi-step reasoning, tool usage, and frequent model switching.
Yes, the SN50 is designed to run multiple models simultaneously using its tiered memory architecture. This allows models to remain resident in memory and be switched quickly with minimal latency. The capability is especially important for agentic workloads that rely on multiple models across task steps, improving responsiveness, utilization, and overall inference efficiency.
The SN50 scales to large models through a combination of memory architecture and distributed deployment. Multiple racks can be interconnected to form inference clusters, enabling support for larger models, higher concurrency, and predictable performance at scale.


.jpg?width=380&height=220&name=Blog%20-%20SN50%20Runs%20the%20Fastest%20MiniMax%20Speeds%20in%20the%20World%20(1).jpg)
.jpg?width=380&height=220&name=Blog%20-%20SambaNova%20Joins%20the%20Genesis%20Mission%20Consortium%20(1).jpg)