Contact

WritingGuide

Running AI agents with a checker

By Rues · Published · Updated

We run most of our work with AI agents: research, code, content, design. The agents are fast, and they are also confidently wrong often enough that we stopped taking "done" at face value. What follows is the setup we use every day, written down so others can copy the parts that fit. It uses Claude Code subagents and plain git, but the roles carry over to other tools.

What are the roles?

There are four roles, and no one plays two of them on the same piece of work.

RoleDoesDoes not
LeadSplits the work, writes briefs, picks the agent and model, mergesWrite the code, design or text itself; approve its own work
AgentDoes one described task on its own branchPush, or touch files outside its brief
CheckerReviews someone else's work, runs tests, returns a verdictEdit files or commit
PersonDecides what goes live, approves spending and visible changesGet skipped

The lead is the main session. It holds the conversation with the person and the overall plan, and it hands everything else out. Keeping the lead from producing work keeps it from grading its own homework.

Each agent is a Markdown file with a short role description and a tool list. Claude Code lets you restrict tools per subagent, and we use that for the checker: it gets read, search and shell tools, and no edit tool. A checker that quietly fixes what it finds ends up approving its own edits.

What goes into a brief?

Agents do not see the conversation, and they do not see each other. The brief is all they know. Every brief has the same seven headings:

  1. Task: what to do, and why. An agent that knows the reason makes better calls when something is unclear.
  2. Your files: the files it may change.
  3. Do not touch: files another agent is working on, other projects, risky areas.
  4. Read first: the documents that hold the rules and context.
  5. Context: what the agent cannot know: what the person said, what was tried and failed, the relevant past decision.
  6. Done means: measurable. "Tests pass", "no unsourced numbers".
  7. Report: a fixed format with the same fields every time: files, summary, tests, open problems, suggestions, and opportunities spotted along the way.

The rule we write the brief by: draw the border, not the path. Say what the agent may touch and when it is finished; leave the method to the agent.

Why does each agent work on its own branch?

An agent that changes files works in its own git worktree, on its own branch. Git worktrees let one repository have several working folders, each on a different branch, so parallel agents do not trip over each other's half-finished files. Agents commit often and never push.

Some git commands are off limits for agents: stash, reset --hard, rebase and push --force. Each of them can hide or destroy work, and an agent running them by mistake is hard to notice afterwards. Files that many tasks touch go to one agent at a time. We learned that one after two parallel agents edited the same file and the merge collided.

How does the checker work?

The checker is a separate agent that did not write the work. It reads the brief and the diff and goes through a fixed list:

  1. Was the task done, and only the task?
  2. Did any file outside the brief change? (git diff --stat)
  3. Do the tests pass, and would they fail without the fix?
  4. Is there any made-up number, date or source?
  5. Were the standing rules followed?
  6. Did a secret end up in code or logs?
  7. Is the text in the right language, and plain?
  8. Did data from one project leak into another?

It answers in a fixed format: a verdict of APPROVE or FIX, then blockers, minor polish, and factual errors such as invented data or an empty test. A blocker must be fixed before the work counts as done.

The checker looks once. If it says FIX, the work goes back to the agent with a short brief, and the lead reads the correction. There is no second checker round. After three rounds of fixes the lead stops and asks the person.

Not everything gets a checker. A small, mechanical change of one file and roughly 30 lines or fewer goes to one agent, and the lead reads the diff itself. Medium work gets one agent and a checker. Large work is split into independent parts with separate files, and each part gets its own checker.

Why isn't a green test proof?

A test that passes with the fix and also passes without it proves nothing. It happens easily: a test that checks too little is the quickest way to green.

So the rule is: after a fix, undo it and run the test again. The test has to fail. If it still passes, it is not testing the fix. Our checker asks this question for every fix.

Green tests can also miss what matters at real scale. In one project, a new component loaded the whole database into memory to show a single result. The tests were green. The checker measured at real data size and found that the page would freeze for 20 seconds. The rule that came out of it: the checker does not approve until it has measured at real data size.

We test the tests in our own site too. Our link test, which fails if any internal link hits a redirect, has a second test that feeds it known bad links and confirms it catches them.

Where does the person come in?

Agents suggest and carry out. A person makes the final call at a few fixed gates:

  • Merging. Only the lead merges, and only after the tests pass and the checker approves. The merge uses git merge --no-ff, which always records a merge commit, so each agent's work stays visible as one unit in the history. The main branch gets changes only through a pull request the person asks for.
  • Visible changes. Before anything visible goes live, the person sees a screenshot at phone width (390 px) and desktop width (1280 px).
  • Money. Before any paid service is turned on, the monthly cost is written down and the person approves it.
  • Secrets. Keys, passwords and tokens never go into files or logs.

The gate exists because agents report success that did not happen. On our home page we list eight such cases, generalized from our own work. In one, the agent said "Change made, operation successful", and when the page was opened, nothing had changed. In another, a sentence was corrected and the replacement was made up too.

What does this cost?

It is slower than letting one agent run free. Every task needs a brief, many need a checker, and some wait for a person. The checker can miss things too. We accept the cost because a wrong "done" costs more to find later than to catch now.

One more habit: before saying "done", check the live result. Open the page and read the line that changed. When a finding will go to someone else, confirm it in a real browser first.

If you are setting up something similar and want to compare notes, write to us.

Sources

Want to talk about this for your brand?

Write about your project and what you want to do. The message goes straight to Rues.

Open the contact form