Blind Checks for Agent Work: Why the Producer Should Not Grade It
When an agent writes its own report and then judges that report, the check is not independent. This guide explains why the agent that did the work should not grade it, and how to run a check that sees only the acceptance criteria and the changed files.
The agent grading its own work
A common agent setup has a task go in and a report come out. That is a solid shape, but it has a hole. If the same agent that did the work also writes the report that says the work is done, the report is the agent grading its own work.
The fix is not a more careful prompt. The fix is a separate agent that checks that report against the criteria. The spawned agent writing its own report is still grading its own work, no matter how honest it tries to be.
The same pattern shows up in a different costume. If an agent keeps going through a fix, break, fix loop, the usual cause is that done was never defined. The agent has no outside test, so it keeps judging its own fix.
Define done before the run
Put the acceptance check in the prompt. That means the test or command that must pass, written down before work starts. Done then becomes something a command can say, not something the agent feels.
A good criterion is one a reviewer can run without the transcript. A task earns its own subagent when it is several tool calls long and its result can be checked that way. If you cannot write a check that stands alone, the task is probably not ready to hand off.
This also keeps the ask honest. Criteria written in advance cannot quietly shrink to match whatever the agent produced.
What blind means in practice
A blind check receives two things: the acceptance criteria and the changed files. It does not receive the producer's summary, its reasoning, or its conclusion.
The reason is anchoring. A reviewer handed the summary starts from the claim that the work is correct, and then looks for reasons it is. A reviewer handed only the criteria and the files has to establish the result for itself. The files are the evidence. The summary is a story about the evidence.
If you build a hand back for the producer, ask for a short typed record: status, files, findings. The record helps the main session. It is not an input to the checker.
The check should be cheap enough to run every time
There is a cost trap on the other side. Having the strongest model verify everything, with the entire chat as context, can cost more than the build itself. People then stop running the check, and the safeguard disappears.
A narrow brief fixes this. When the verifier gets only the criteria and the changed files, the input is small and the job is well bounded. That keeps the check cheap enough to run on every task, which is the only way it protects you every time.
The same logic applies to long reads and test runs in general. Where a subagent pays off is in taking work with a lot of raw output, so the main context gets a short hand back instead of the raw logs. A check is one more case of exactly that.
Checks come after the work
Order matters. A check that runs alongside the work, or before it is finished, can only check a moving target. Run it as its own step after the work is done, then read the verdict.
Wave based dispatch makes this natural. Work in waves: research first, then produce, then integrate, then check. The check wave is last by construction, so nothing can be graded before it exists.
When a check fails, the failure goes back to the producing role with the reason attached. The checker does not repair the work itself, because then it would be grading its own repair.
Keep the producer away from the tests
Loosening the tests is the classic way an agent gets to a green result. Two controls stop it without anyone needing to flag it.
- Keep test files out of the worker's write scope.
- Have a separate checker run the original suite, not a version the worker has touched.
With both in place, the only way to pass is for the code to satisfy the tests that were there at the start.
Write the role, not just the task
With several agents coordinating, structure lives in what you write per role: what it does, which files it may touch, what done means, and who checks it. Writing that before the run keeps the coordination out of the chat.
The last item, who checks it, is the one people leave out. Name the checker in the plan, give it its own brief, and make sure the brief carries nothing from the producer. Then the log can show, for any change, who made it and who independently looked at it.
Escalate on failure, not on feeling
A failed check is also the best trigger for moving a task to a stronger model. It is concrete, it is recorded, and it comes with a reason. Escalating on a vague sense that a task is hard sends most of the work back to the expensive model, so use the check result as the signal instead.
Whether you do all this by hand or with a tool, keep the principle: the producer produces, a different role checks, and the check sees only criteria and files. The next section shows one way to make that a role in the plan.
Making it a plan with crews
crews is a Claude Code plugin that plans subagents before they run: you list the roles, each gets a fixed model and effort, they run in waves, and a hook blocks agents that were not planned. Check roles are blind, so they get criteria and file paths and never the producer's conclusion.
claude plugin marketplace add mdalexandre/crews claude plugin install crews@crews
Linux and macOS, needs uv. crews is an independent project, not affiliated with Anthropic.
View crews on GitHub