Claude Code Subagents Inherit the Main Model: How to Find and Pin Them

If a subagent definition has no model line, it runs on whatever model the main session uses. That means the plan can say one thing, the worker can run another, and routine file reading can quietly land on Opus. This guide covers how to find those agents, how to confirm what actually ran, and how to pin a model per role.

By Mario AlexandreMCP & Agents

The gap that a scan misses

A common first step when usage runs hot is to scan your agent files for the ones that pin an expensive model and change them. That step is worth doing, but it leaves a gap. The files you most need to find are the ones that say nothing about a model at all.

The reason is simple. An agent definition without a pinned model inherits the parent. If your main session is on Opus, every such agent runs on Opus too. Nothing in the file warns you. It reads as neutral, and it behaves as the most expensive choice you have.

So the scan has to be inverted. Do not only list the agents that pin a different model. List the agents that have no model line, because those are the ones that follow the main session wherever it goes this week.

This matters most for plain work: reading files, searching, running tests. In my experience those jobs finish fine on Sonnet at medium effort. When they land on Opus instead, they are paid for at Opus rates and they add nothing the cheaper model would not have produced.

Step one: list the agent files with no model line

Start with the cheapest check, which is a file search. You want every agent file that has no model: line.

grep -L "^model:" <your agent files>

Read the output as a to do list. Each file it prints is an agent whose model is decided by the main session at the time it runs. Decide for each one whether that is what you want. For a planner or a hard debugging role, inheriting a strong main model may be exactly right. For a file reader, a search role, or a test runner, it usually is not.

Keep this list. You will want it again after you pin models, because the same grep should come back empty for every role that does routine work.

Step two: confirm the model a worker actually ran on

A file check tells you what a definition says. It does not tell you what ran. The only way to know is to look at the run itself.

Open a subagent transcript and confirm the model it actually ran on. An agent definition without a pinned model inherits the parent, so an instruction can hold in the plan and still not reach the worker. You can have a clear intent, a sensible plan, and a worker on the wrong model, and the transcript is where that shows.

The same gap exists one level up from any status display. A header or status line that shows the model for each message is a good guard, but a subagent can run on a different model than the one shown. Logging the model each worker actually ran on closes that gap, and it turns a vague worry into a line you can read.

Step three: count workers and models per session

When the same workload suddenly costs more, look at the subagent level before you blame the main session. Count how many subagents each session spawned, and note which model each one ran on.

Ones with no model set run on Opus when the main session does. A few extra workers can account for the jump. A single prompt that fans out can also drain a window fast, because each worker pays its own way through the same limit.

A run with fifteen subagents is where the weekly limit tends to go, and usually most of those workers are reading files and running tests. That is the typical shape. The strong model is useful for the plan and for the hard calls. The bulk of the worker count is ordinary reading and checking, which a cheaper seat handles well.

Step four: pin a model per role

Once you know which agents inherit, give each role an explicit model. A short rule set is enough to start.

  • Planning and hard bugs: keep the strong model, at high effort.
  • File reading, searching, test runs, routine edits: pin to Sonnet, at medium effort.
  • Open ended debugging: be careful. It can take enough retries on a cheaper model to end up costing more, so split by task shape rather than by habit.

Pinning is the first fix to try when Opus on a medium setting still burns the limit faster than expected. Plain file reading and searching at Opus rates is a quiet cost, and pinning those agents to Sonnet removes it without touching the part of the work that needs the strong model.

Per task cost is the right comparison here, and it depends on the task. Bounded edits and file reading usually finish without trouble on Sonnet. A task that needs judgment is the one to keep on the stronger model.

Pin the model on headless calls too

Subagent files are one place where a default can drift. Scripts that start headless Claude calls are another. If those calls do not name a model, the children follow whatever the default is this week.

Pass an explicit model on every headless claude call. Then log the model each child reports back. A silent change in the default will show up in the log as a different model name, instead of showing up later as a larger bill.

The pattern is the same at each level: say the model out loud where the work starts, and record the model that actually ran. Declared and observed should match, and when they do not, you want to see it the same day.

Escalate on a concrete trigger

Pinning cheap by default only works if you also say when to go up. Cheap by default, escalate on complexity is a good rule, but the trigger has to be concrete. A failed check is a good trigger. So is a named risk, such as a change that touches shared state.

A vague sense that a task is hard is not a good trigger. Escalating on a feeling sends most of the work back to the expensive model, and you are back where you started. Write down what moves a role up, so the choice is made before the run and not in the middle of it.

Making it a plan with crews

crews is a Claude Code plugin that plans subagents before they run: you list the roles, each gets a fixed model and effort, they run in waves, and a hook blocks agents that were not planned. Check roles are blind, so they get criteria and file paths and never the producer's conclusion.

claude plugin marketplace add mdalexandre/crews
claude plugin install crews@crews

Linux and macOS, needs uv. crews is an independent project, not affiliated with Anthropic.

View crews on GitHub