OVHcloud Brings SambaNova to Premium AI Endpoints Across Europe
OVHcloud serves more models in less space with record-breaking inference speeds, powered by SambaNova's full-stack technology.
TL;DR
- OVHcloud is bringing SambaNova into its Premium AI Endpoints, delivering Europe's most advanced AI inference platform.
- The endpoints run on SambaRack SN40-16, powered by SambaNova's Reconfigurable Dataflow Units (RDUs), at an average of 10 kW per rack.
- SambaNova's three-tier memory architecture lets OVHcloud serve more models with less hardware, with multiple models per SambaRack hot-swapped at runtime with very low switching times.
- RDUs deliver unprecedented throughput-per-watt and roughly 4x better energy efficiency over traditional GPUs.
- OVHcloud Premium AI Endpoints are available now, serving open-source models including OpenAI's gpt-oss and DeepSeek.
Background
Inference has overtaken training as the dominant AI workload, and it draws power continuously rather than in bounded runs. For an established European cloud provider, adding premium inference capacity is a question of throughput, energy efficiency, and how many models a given footprint can serve, all on mission-critical infrastructure.
OVHcloud selected SambaNova to power its Premium AI Endpoints, bringing record-breaking inference speeds to European customers on hardware built for density and efficiency.
Challenge: More Models, Less Space, Mission-Critical Uptime
OVHcloud's requirements were shaped by scale.
- Throughput per watt. Premium inference has to deliver speed without a runaway power bill.
- Model density. Customers expect a broad catalog, which is uneconomical if every model needs its own hardware.
- Mission-critical reliability. The endpoints serve generative AI workloads that cannot fail.
- A credible open-source catalog. Developers expect current frontier open weights.
Solution: SambaRack SN40-16 and Three-Tier Memory
OVHcloud standardized on the SambaRack SN40-16, built on SambaNova's RDUs.
Unprecedented Throughput Per Watt
SambaNova's RDUs enable OVHcloud to deliver unprecedented throughput-per-watt, with roughly 4x better energy efficiency over traditional GPUs. Each SambaRack SN40-16 runs large models at 10 kW average power.
Efficient Inference Through Three-Tier Memory
SambaNova's three-tier memory architecture spans large-capacity DDR, high-bandwidth HBM, and on-chip SRAM, keeping model weights and data close to compute. That hierarchy cuts the data movement that limits GPU-based inference, so OVHcloud sustains high throughput on large models while holding energy consumption down. The memory architecture is what turns raw RDU performance into efficient, mission-critical inference.
Record-Breaking Inference Speeds
OVHcloud's Premium AI Endpoints, powered by SambaNova's full-stack technology, deliver record-breaking inference speeds on mission-critical generative AI workloads, serving open-source models including OpenAI's gpt-oss and DeepSeek.
Why OVHcloud Chose SambaNova
- Throughput per watt. RDUs deliver unprecedented throughput-per-watt and roughly 4x better energy efficiency over traditional GPUs.
- Model density. Three-tier memory lets OVHcloud serve more models with less hardware, hot-swapped at runtime.
- Efficiency at 10 kW. Each SambaRack SN40-16 runs large models at 10 kW average power.
- Open-source breadth. Current open-weight models, including gpt-oss and DeepSeek.
What Comes Next
OVHcloud Premium AI Endpoints launched in 2026. This is the scale-cloud proof point: An established European provider selecting RDUs for throughput-per-watt and model density on mission-critical generative AI. Read more on sovereign AI and national autonomy, or talk to our team about a deployment in your region.
10 kW average power per SambaRack SN40-16
Roughly 4x better energy efficiency over traditional GPUs
Multiple models per rack, hot-swapped at runtime
Premium AI Endpoints launched in 2026
“SambaNova provides the raw power and efficiency we demand for premium AI. Our partnership lets customers deploy more models in less space, achieving enterprise-grade inference no other provider can match.” — Octave Klaba, CEO, OVHcloud
FAQs
Premium AI Endpoints, powered by SambaNova's full-stack technology, deliver Europe's most advanced AI inference platform with record-breaking inference speeds.
SambaNova's three-tier memory architecture lets multiple models run per SambaRack and hot-swap at runtime with very low switching times, so OVHcloud offers a wider catalog from a smaller footprint.
The SambaRack SN40-16, built on SambaNova's Reconfigurable Dataflow Units, runs large models at 10 kW average power with roughly 4x better energy efficiency over traditional GPUs.


.jpg?width=380&height=220&name=MiniMax2.7Organic%20(1).jpg)
