Ricoh runs Japanese custom AI models 10× faster on SambaCloud

Enabling agentic AI for Japanese businesses, from tens to over 700 tokens per second

ricoh_runs

Ricoh deployed SambaNova's SambaCloud to power its custom Japanese AI models, achieving 10x faster inference speeds and over 700 tokens per second on 70B-class models, reducing complex agentic workflow times from one minute to ten seconds.

Ricoh, a global Japanese digital services and electronics company, is renowned for its imaging products and services for businesses and consumers.  As one of the world’s largest manufacturers of cameras, printers, photocopiers, and projectors, Ricoh has also established itself as a leader in document management systems offered as SaaS solutions. Building on this foundation, Ricoh has been actively expanding into cutting-edge digital technologies, including artificial intelligence (AI), to drive innovation and growth.

What they do:

Beyond its traditional electronics and imaging business, Ricoh has emerged as a key player in Japan's AI ecosystem. The company has developed AI models and services tailored specifically for Japanese businesses, addressing unique linguistic and cultural nuances.

Using open-weight models like Llama, Qwen, and Gemma, Ricoh has developed specialized models for handling Japanese business documents and industry-specific models. One standout offering is RICOH Digital Buddy, a generative AI agent that leverages internal documents and organizational knowledge to answer questions and execute tasks, streamlining workflows for businesses.

Internally, Ricoh has embraced AI to enhance operational efficiency. By adopting no-code platforms like Dify, the company empowers non-engineers to build and deploy their own AI workflows, democratizing AI usage across the organization.

Challenge:

Ricoh's core challenge was that existing GPU infrastructure could not run 70B-class models efficiently enough for production agentic workloads, achieving only tens of tokens per second.

Ricoh’s success in developing a diverse range of custom AI models for Japanese businesses has brought new challenges. Bringing these models into production at scale requires an environment that hosts them efficiently from a model-provider standpoint and ensures fast inference from an end-user standpoint. This requirement is becoming increasingly pressing with the rise of modern agentic workflows, which involve calls across multiple models, not just a single call to one model.

Existing infrastructure solutions fell short of these requirements. Lightweight, low-cost GPU environments struggled to run 70B-class models efficiently, achieving only tens of tokens per second. On the other hand, high-end GPU environments raise concerns about scalability and cost-effectiveness, making them less viable for Ricoh’s long-term goals.

Solution:

Ricoh turned to SambaCloud, SambaNova's fully-managed cloud inference service, to overcome these challenges. Powered by the SambaNova Reconfigurable Dataflow Unit (RDU), SambaCloud is designed to efficiently run a wide variety of open-weight models, delivering both speed and scalability.

By adopting SambaCloud, Ricoh achieved:

  • 10× faster speeds compared to their existing infrastructure, with over 700 tokens per second on their 70B-class models.
  • High accuracy and performance, ensuring their fine-tuned models optimized for Japanese business contexts maintained their effectiveness.
  • Scalability and reliability, enabling Ricoh to support modern agentic workflows and meet the growing demands of their AI-driven solutions.

Why Ricoh Chose SambaCloud:

Ricoh’s decision to partner with SambaNova was driven by several key factors:

  1. Unmatched Performance: SambaCloud’s ability to deliver 10× faster inference speeds directly addressed Ricoh’s need for high-performance infrastructure.
  2. Cost-Effectiveness: SambaCloud provided a scalable solution that balanced performance with cost, making it a sustainable choice for Ricoh’s expanding AI initiatives.
  3. Specialized Support for Open-Weight Models: SambaCloud’s compatibility with a wide range of open-weight models, including those fine-tuned by Ricoh, ensured seamless integration and deployment.
  4. Focus on Japanese Business Needs: SambaCloud’s infrastructure preserved the accuracy and cultural relevance of Ricoh’s models, which are tailored for Japanese businesses.

Results:

With SambaCloud, Ricoh has successfully scaled its AI operations, enabling faster and more efficient deployment of their custom models. This partnership has empowered Ricoh to continue innovating in the AI space, delivering cutting-edge solutions to Japanese businesses while maintaining their leadership in the industry.

10x

faster

700+

tokens per second

“With SambaNova running 5 to 10 times faster, even complex agentic workflows that would otherwise take a minute finish in 10 seconds. We think that brings significant business value.”

 

— Gakushi Miyara, AI Service Business Division

Ricoh Company, Ltd.

 


Mr. Miyara, AI Engineer at Ricoh, shares his experience deploying custom AI models on SambaCloud.

 


Find out the business value of SambaCloud to Ricoh.

FAQs

What is SambaCloud and how does Ricoh use it?

SambaCloud is SambaNova's fully managed cloud inference service, powered by the Reconfigurable Dataflow Unit (RDU). Ricoh uses it to host and serve custom AI models fine-tuned for Japanese business documents and industry-specific workflows.

How much faster is SambaCloud than standard GPU infrastructure?

SambaCloud delivers 10x faster inference than Ricoh's existing GPU infrastructure, achieving over 700 tokens per second on 70B-class models compared to tens of tokens per second previously.

Does SambaCloud support fine-tuned open-weight models?

Yes. SambaCloud is compatible with a wide range of open-weight models including Llama, Qwen, and Gemma, and supports fine-tuned variants, preserving accuracy and cultural relevance for Ricoh's Japanese business use cases.

How does SambaCloud handle agentic AI workflows? 

SambaCloud's speed enables complex agentic workflows involving multiple model calls to complete in around ten seconds, compared to approximately one minute on previous infrastructure.

Back to top

It’s all about you

SambaNova Expands Deployment with SoftBank Corp. to Offer Fast AI Inference Across APAC

SambaNova Expands Deployment with SoftBank Corp. to Offer Fast AI Inference Across APAC

March 5, 2025
Qwen3 Is Here - Now Live on SambaNova Cloud

Qwen3 Is Here - Now Live on SambaNova Cloud

May 2, 2025
SambaNova Partners with Meta to Deliver Lightning Fast Inference on Llama 4

SambaNova Partners with Meta to Deliver Lightning Fast Inference on Llama 4

April 7, 2025