In this short DevTalks clip, Justin Woo shows something easy to overlook when building agents: the model behind each agent is a setting you can change, and changing it takes seconds. Models vary in size and in cost, so the right choice depends on what that particular agent is being asked to do.
TL;DR
- The model is set per agent, not per project. In the demo every agent is running DeepSeek V3.1, and one is switched to a different model by selecting it and saving.
- Smaller can be the better call. DeepSeek V3.1 is a large reasoning model; Llama 4 Scout is much smaller and much more cost effective, and SambaNova hosts models across that range at different price points.
- Test the change without rerunning everything. Replaying a single task shows the effect of the new model on its own, and on a smaller model it finishes almost immediately.
Swapping the model behind an agent
In the demo, each agent in the crew is running DeepSeek V3.1. Justin picks one, opens the model list, and switches it. SambaNova hosts a range of models, some larger and some smaller, with different pricing attached. DeepSeek V3.1 is a large reasoning model, so the swap here is to something lighter: Llama 4 Maverick is an option, and he settles on Llama 4 Scout, which is much smaller and much more cost effective.
The change is just select and save, and the agent uses the new model from then on. The useful part is what comes next: rather than rerunning the whole crew, he replays that single task to see the effect in isolation. On the smaller model it completes so quickly there is barely time to narrate it.
How to decide which model fits a task
Start with published benchmarks
The recommendation is to use something like Artificial Analysis to work out which model suits a particular agent’s task. Models differ in what they are good at: instruction following, coding, scientific reasoning, and agentic tool use are all measured separately, and an agent that is strong on one is not automatically strong on another.
Then check your own domain
There are a lot of benchmarks available, so the second step is to look at the one that matches your field rather than stopping at the general leaderboard. Finance and healthcare are the examples given. The benchmark that reflects your actual work is the one worth optimizing against.
FAQs
Can different agents in the same crew use different models?
Yes. The model is a per-agent setting. In the demo all the agents start on DeepSeek V3.1, and one is moved to a different model by selecting it from the list and saving, without touching the others.
Why use a smaller model instead of a large reasoning model?
Cost. DeepSeek V3.1 is a large reasoning model, while Llama 4 Scout is much smaller and much more cost effective. SambaNova hosts models across that range with different pricing, so matching model size to what the task actually needs keeps spend down.
How do you work out which model is best for an agent task?
Start with a benchmark aggregator such as Artificial Analysis and compare models on the capability that matters for that agent, whether that is instruction following, coding, scientific reasoning or agentic tool use. Then look at benchmarks specific to your own domain, such as finance or healthcare, rather than relying on general rankings alone.
