Heads up: I’m aiming to open source the orchestrator. Not there yet, but I’m writing these posts on the way.
Edit: While it’s been exciting working on the AI agent “factory”, we’ve reached the end of the line with the approach as described. Claude Code has gotten better and better in a short amount of time that it’s inevitable for orchestration to be best directed by the providers themselves. So it no longer makes sense to open source the orchestrator which I am retiring. However (!), I am continuing my quest for better tooling for making products with AI. So I will keep writing about AI productivity.
Part IV was about an end-to-end test harness so I could trust the factory to run without me at the keyboard. Part V is about a problem I’d been consciously ignoring the whole time: the factory was marking its own homework.
Every agent worker ran on the same family of model, Claude. The worker that drafts the spec, writes the code, reviews the pull request are Claude Sonnet or Opus. It felt easier to start with only Claude but the same family of model writing code and grading it is a conflict of interest.
The AI orchestrator, in summary
Workers in the factory can only “propose” outputs. The outputs are owned by the AI agent orchestrator. Those are the rules.
If you’ve not read the earlier posts in this series, the AI agent orchestrator we’ve been building picks up GitHub issues, spins up worktrees and gets AI workers to do the work. Research a spec, write code, review the PR. The workers do non-deterministic output generation (code, writing, etc). The orchestrator facilitates deterministic routing and execution of work.
A worker can’t open a pull request, push a branch, move a card on the board or close an issue. All it does at the end of a run is write a small file called a “verdict” and the spec, diff, etc. The orchestrator reads that verdict file and performs the changes.
In the earlier posts I described this as a safety mechanism. Dangerous capabilities were taken away from the agent worker rather than it being instructed not to use them, because an LLM that’s been told “don’t push to main branch” will, eventually, push to main. So the orchestration has safety guardrails built in. The worker proposes and the orchestrator has the final say. A worker that hallucinates “close fifty issues” can’t close them.
What I didn’t appreciate until recently is that this safety rule also made model selection extensible.
The conflict of interest
I knew this was a problem but kept ignoring it. The same family of model was writing code and reviewing it. Even when those are different Claude models (Sonnet writing, Opus reviewing), they’re still sharing assumptions, blind spots and stylistic habits. The reviews looked thorough but didn’t benefit from adversarial review.
Sibling models from the same vendor can lead to the reviewer going easy on the author. This is a known effect: a NeurIPS 2024 study found an LLM evaluator has self-bias — it scores its own outputs higher than it scores others’, even when humans judge them equal. The bias doesn’t stop there: a separate study found family bias too, the same effect across siblings (Claude 3.5 Sonnet rated other Claude models higher).
While those papers measure scoring and not bug-catching, shared training and blind spots can wave through a bug the author is missing.
The fix is to use a different vendor’s model. A reviewer that didn’t write the code and doesn’t share the author’s (i.e. model’s) blind spots.
Like a friend once said to me, “if you want to check a SQL query’s correctness, have someone else review it.” Test with different eyes on the same problem.
I decided to route reviews through Cursor’s CLI instead of Claude, running Cursor’s own model (Composer).
Why the change was almost free
Changing the PR reviewer’s model was surprisingly easy.
The review worker is a subprocess. The orchestrator launches a command-line AI agent in a worktree (claude -p), hands it a prompt and waits for it to write its verdict file. To review with a different vendor, a second adapter was needed that launches Cursor’s CLI (cursor-agent), which has a headless mode that prints structured output you can parse. It has the same shape as the Claude CLI already in place.
I was already using Cursor for a different reason: when a problem needs more hand holding than unattended agents can give it, I sit at the keyboard in Cursor and steer. I had an annual subscription I wasn’t using as much as I should and their agent headless CLI made it a viable reviewer option.
The reason this was a smooth change is because the worker only produces the verdict and result. The worker was already decoupled from making the changes. The orchestrator reads the verdict file, runs the validation, posts the review and logs what the run cost. The model was already a configurable variable, even though I never used anything other than Claude.
Cursor is the default reviewer now; a label on the issue can override the vendor. The orchestrator reads the label and dispatches the appropriate adapter and everything else is identical. The benefit of a clean boundary is that I built workers with safety in mind and it ended up making the system more extensible.
Least privilege, no credentials
The provider interchangeability also means the Cursor worker doesn’t get a GitHub token.
A review worker needs to view the pull request’s diff, description, comments. One way is to hand it a token and let it fetch the PR. But the moment you hand a worker a token, you’ve given it keys to cause potential damage. Safety depends on the token’s scopes being right.
The worker doesn’t get keys. Before it starts, the orchestrator (which holds credentials) fetches the diff, writes it into the worktree in a plain file and checks out the PR’s branch. The worker reads the diff, reasons about it, writes its review. The worker doesn’t access GitHub. It couldn’t write a rogue comment even if it tried because it has no idea GitHub exists. The worker is decoupled from GitHub.
The Claude worker also has a hook that intercepts forbidden commands. Cursor’s CLI doesn’t have the hooks I have for Claude, so we don’t give Cursor the keys or tokens. The typed verdict file works the same way. The worker can’t post the review itself, only write a file the orchestrator then acts on. We’ve designed the mistake out: poka yoke.
On a separate note, Cursor runs on my existing Cursor subscription, a different bill from my Claude usage, so reviews go on a different budget to Claude’s.
What I haven’t done yet
This is a very confined, focused change. One workflow changed (PR reviewer), replacing one model for a different model. Other workers still run on Claude.
Nothing auto-merges (yet). A different model reviewing the code in theory is a better approach but not a replacement for a human reading complex PRs. The factory has always required a human to approve before anything ships and that hasn’t changed.
One rough edge: the reviewer still posts under my own GitHub identity, not its own. In an ideal world it would have a separate account.
It’s early days, so I’m not going to claim it catches things Claude wouldn’t. Maybe that shows up as the reviews pile up and I’ll learn as we go. The first runs show that reviews follow the repo’s conventions, findings are attributed to a file, line and fix.
The model can be swapped
I drew a line between the non-deterministic part (LLM workers) and the part that touches real code and files (orchestrator). The side effect is an ideal environment for model agnosticism, safety and extensibility.
I can now flip a label on an issue to choose which model reviews it. The factory grows!
If you’re looking to use AI structurally in your workflow, I’d love to hear from you! I write more like this at blog.mariohayashi.com, and feel free to follow me on X: @logicalicy.



