Unattended Agent Run Checklist: What to Set Before You Leave
Overnight and unattended runs fail in a predictable way: nothing stops the fan out while you sleep. This checklist covers what to set before you leave, so the run is predictable when you are not there to watch it.
Why unattended runs go wrong
A run you are watching has a natural brake: you. You see a second research agent start, you see an audit begin, and you step in. A run you started before bed has no such brake.
A setup of six research agents plus four audits, each free to spawn more, is how a cheap model ends up expensive. The count is not the only problem. The real issue is that nothing in the setup says stop.
Earned complexity applies to worker count too. Start with one, and add another only when the main context is actually diluting. Once runs go unattended, a hard cap on workers matters more than which model each one gets. Use the checklist below in that spirit: every item is a limit you set in advance.
Item one: a hard cap on workers, and no nested spawning
Set a hard cap on workers per run. Then add the second half of the rule: workers never spawn their own. Together those two lines made my unattended runs predictable.
If you also run a review loop that keeps spinning up new agents, put a hard cap per round. Without one, the loop can add workers faster than it closes tasks.
A queue helps here. A first in, first out queue is the part people skip, and it is what makes a small cap usable overnight. With a cap of two workers and a queue behind it, work still finishes, just in order.
Item two: a model pinned per subagent
Every subagent gets an explicit model. An agent with no model set runs on whatever the main session uses, so file reading and test runs can land on the most expensive model without anyone choosing that.
Decide the split before the run. Planning and the hard calls go on the strong model. Reads, edits, and test runs go on a cheaper one. Deciding it per role turns the bill into something you set instead of something you discover.
Write the model next to the role, not in your head. A reviewer, or you the next morning, should be able to compare what ran with what was planned.
Item three: a turn limit and a minimal tool list per worker
One stuck worker can eat the budget. A turn cap per subagent stops a bad loop from running all night. A tool list limited to what the job needs removes whole classes of mistakes, because a worker that cannot edit a file cannot edit the wrong one.
State both next to the model in the plan, so they can be checked too. A plan that names the model but leaves turns and tools open is only part of a plan.
A reader role needs read tools. A test runner needs the command it runs. Neither needs the ability to push, and the list should say so by leaving it out.
Item four: separate worktrees and a worker report
With several agents in parallel, the first thing to rule out is file collisions. Give each agent its own git worktree, or at least its own write scope, so agents do not edit the same files at once. If you see index lock collisions, a separate worktree per worker helps there too, and only the merge step touches the main checkout.
Add one habit that saves a search later. Have each worker end by stating its worktree path and whether it committed. A fresh session then reads a short list instead of hunting through every worktree.
Naming helps as well. When each agent is named for its job and has its own write scope before it runs, and its hand back lists the files it changed, the log answers who touched what.
Item five: a check that sees only the criteria and the files
Do not let the worker grade itself. Put the check in its own step after the work, and give it only the acceptance criteria and the changed files. It does not need the whole chat and it should not see the producer's summary.
There is a cost benefit too. Having the strongest model verify everything with the whole conversation in front of it can cost more than the build itself. A narrow brief keeps the check cheap enough to run on every task.
Keep test files out of the worker's write scope, and have the checker run the original suite. Loosening the tests is the classic failure, and this stops it before anyone needs to flag it.
Item six: a fixed return shape for hand backs
Free form summaries drift long. The main session pays for every extra line again on each later turn, so a long hand back is a cost that keeps growing.
Give each worker a fixed shape: status, files changed, and one line per finding. For the orchestrator itself, ask it to end each run with a fixed report of status, files changed, and open decisions, and keep the step by step in its log instead of the chat.
The shape also makes the morning review fast. You read the same few fields from every worker, and a missing field is itself a signal.
Item seven: one session per task and few pushes
Two smaller habits belong on the list. Start a new session for each task instead of one long one, so old context is not carried and paid for again.
Push once per finished task, and run only the affected tests on intermediate commits. Running the full suite on every push is a multiplier, and agents can push far more often than a person would.
Write it as a plan, not as a hope
Every item above is a decision you can make before the run. The point of a checklist is that it moves those decisions out of the moment. Write the roles, what each may touch, what done means, and who checks it. Then dispatch in waves, so coordination stays out of the chat.
If you want that written once and enforced, the next section describes one tool built for it.
Making it a plan with crews
crews is a Claude Code plugin that plans subagents before they run: you list the roles, each gets a fixed model and effort, they run in waves, and a hook blocks agents that were not planned. Check roles are blind, so they get criteria and file paths and never the producer's conclusion.
claude plugin marketplace add mdalexandre/crews claude plugin install crews@crews
Linux and macOS, needs uv. crews is an independent project, not affiliated with Anthropic.
View crews on GitHub