Matthew Boston

Two Code Review Loops I Handed to Agents

July 28, 2026

Loops are all the rage right now, so here are two I’ve been running. Both live inside code review, and code review is mostly waiting. When I open a pull request, I wait for comments. When I leave a comment, I wait for the fix. I built an agent for each side of that wait, and both run in the background while I work on something else.

Code review is two loops

From the author’s side, a pull request goes like this: open it, wait for comments, read each one, decide which ones matter, fix those, push, and wait again. From the reviewer’s side, you leave a comment, wait for the author to respond, then go back and check whether the change addresses what you asked. Both loops repeat until somebody approves.

Look at where the time goes. The judgment is the short part. Deciding whether a comment deserves a change takes seconds once you’re looking at it, and checking a fix takes a minute or two. The rest is watching threads and loading a PR back into your head after you stopped thinking about it hours ago.

Watching a thread is work a machine can do, so I handed it to one.

The author loop

The first agent watches my open PRs. When a comment lands, it triages the comment against our skills and context, then applies the fixes.

Triage is what makes this worth having. A review comment might be a real bug the reviewer caught. It might be a style preference the team’s conventions already settle. It might be a question that needs an answer and no code change at all. Or it might suggest something the team already decided against, for reasons written down somewhere the reviewer didn’t look. An agent that applied every comment literally would churn the diff for every passing opinion and quietly undo decisions that had reasons behind them.

The skills and context give the triage something to judge against. The same SKILL.md files that teach an agent how to build a feature can tell it whether a reviewer’s suggestion matches how the codebase already does things. Without that context, an agent has no basis for weighing a comment, and the only thing left for it to do is whatever the comment says.

The reviewer loop

The second agent watches the PRs I’ve commented on and checks whether the resolution holds.

A resolved thread means someone clicked a button. It doesn’t tell you whether the new commit fixed the problem. The author may have fixed the line the comment pointed at and missed the same mistake two functions down. They may have changed something that matches the wording of the comment and misses its point. They may have replied with a reason they disagree, which is a legitimate answer, and one the reviewer should read before the PR merges.

Checking that is the same discipline as reviewing the outcome instead of the output. The question is whether the change did what the comment asked for. An agent can hold the comment and the follow-up commit side by side and answer that, without me reopening the PR and rebuilding the context from scratch.

What makes a background loop safe

Running agents in tmux panes I can see is one thing. An agent that works while nobody’s watching needs tighter boundaries, and any loop like these depends on a few of them.

  • It needs a stop condition. An author-side agent and a reviewer-side agent on the same PR could pass a thread back and forth forever if nothing tells them when to quit.
  • Its fixes go through the same checks as everyone else’s. A change applied in the background still has to pass lint, types, and tests, which only works if that feedback loop is fast enough to run on every push.
  • It knows which comments to leave alone. A comment that questions the design of the feature needs a human decision. The agent should flag it and move on.
  • Its changes stay small. A fix for review feedback should be about the size of the feedback. If a one-line comment produces a forty-line diff, the triage went wrong somewhere.

None of these rules are new. They’re what you’d tell a junior engineer the first time they handled review feedback without someone looking over their shoulder.

Less time watching threads

Code review needs my judgment at a few specific moments: when a comment calls for a real decision, or when a fix doesn’t hold up. Everything between those moments is waiting, and the agents do the waiting now.

Both run in the background. I spend less time watching threads and more time building.