OpenRouter Uses SambaCloud to Deliver High-Speed LLM Performance

Developer focused company meets the performance demands of real-time use cases

2025-05-19_OpenRouter+SN_v1.2

Challenge:

OpenRouter is the largest marketplace for LLM inference. They have hundreds of LLMs available across dozens of different cloud providers. In cases where developers and enterprises need inference for their product or use case, they build a single integration to OpenRouter and they can access all the different models that they may need through a single interface and a single billing relationship.

Some developers require consistent, high speed inference for their use cases. In these cases total generation times are dominated by the overall throughput and not just time to first token.

As a developer-first platform, OpenRouter found a wide variety of different use cases that need the throughput that SambaNova offers. In particular, this high performance is critical to interactive uses cases where a user may be in a chat mode and an instant response is required.

Solution:

OpenRouter offers a variety of open source models via SambaCloud. As a high-speed inference provider, SambaNova hits the sweet spot of price and performance for a number of OpenRouter customers who require reliable inference and high throughput.

OpenRouter uses SambaCloud. Powered by SambaNova RDU, the SN40L, SambaCloud delivers the fastest inference on the largest and best open source models., including the fastest performance available for DeepSeek R1 671B. Meta Llama 4 Maverick, and more.

100s

LLMs available across dozens of providers

>6

Average number of models companies use

50+

Providers on their platform

“SambaNova, as a high throughput provider, has an important role in the market for use 
cases where they care about the total time to last token for larger prompts."


— Chris Clark, OpenRouter Chief Operating Officer



In the above video, Chris Clark, Chief Operating Officer at OpenRouter discusses how OpenRouter leverages the performance of SambaNova to meet customer demands.

FAQs

What is SambaCloud and how does Ricoh use it?

SambaCloud is SambaNova's fully managed cloud inference service, powered by the Reconfigurable Dataflow Unit (RDU). Ricoh uses it to host and serve custom AI models fine-tuned for Japanese business documents and industry-specific workflows.

How much faster is SambaCloud than standard GPU infrastructure?

SambaCloud delivers 10x faster inference than Ricoh's existing GPU infrastructure, achieving over 700 tokens per second on 70B-class models compared to tens of tokens per second previously.

Does SambaCloud support fine-tuned open-weight models?

Yes. SambaCloud is compatible with a wide range of open-weight models including Llama, Qwen, and Gemma, and supports fine-tuned variants, preserving accuracy and cultural relevance for Ricoh's Japanese business use cases.

How does SambaCloud handle agentic AI workflows? 

SambaCloud's speed enables complex agentic workflows involving multiple model calls to complete in around ten seconds, compared to approximately one minute on previous infrastructure.

Back to top

Related resources

SambaNova Expands Deployment with SoftBank Corp. to Offer Fast AI Inference Across APAC

SambaNova Expands Deployment with SoftBank Corp. to Offer Fast AI Inference Across APAC

March 5, 2025
Qwen3 Is Here - Now Live on SambaNova Cloud

Qwen3 Is Here - Now Live on SambaNova Cloud

May 2, 2025
SambaNova Partners with Meta to Deliver Lightning Fast Inference on Llama 4

SambaNova Partners with Meta to Deliver Lightning Fast Inference on Llama 4

April 7, 2025