Tech Analysis
What a Hospital for Coding Agents Teaches About Review
October 2026
A database company let AI coding agents write a large share of its migration tooling for five months, then published what happened. The line count is eye-catching, but the more useful part is what stood between the agents and the main branch. This post covers that design, where it struggled, and a version a one-person project could realistically use.
The setup
Cockroach Labs, which makes the CockroachDB database, built a pipeline called MOLT Sinai and modelled it on a teaching hospital. Each GitHub issue is a patient, merging is discharge, and the humans in charge are the chiefs of medicine. Agents take named roles with their own instructions. A Fellow investigates and fixes. A Review Attending, a separate agent told to hunt for faults rather than fix them, critiques the work. A Discharge Nurse makes the final check before merge.
Their reasoning was about priorities. A subtle bug in a migration tool could corrupt a customer's data, so the question they cared about was how many merged changes they would be embarrassed by, not how many arrived per hour.
The rules that did the work
- A plan before code, and a review before the plan is used. The agent must reproduce the problem, diagnose it and post a plan listing files, tests and risks. A second agent approves or rejects it first.
- No improvising. If the scope changes mid-fix, the agent stops and goes back for a revised plan. The instruction in its file is just two words: "Don't improvise."
- Tests are not negotiable. An agent may not weaken, skip or edit a test to make it pass. Reviews start by switching the change off and confirming the new tests fail.
- Stuck means hand off, not try harder. A stuck agent writes a structured handoff for a senior agent, which must restate what it understood before starting.
- The last gate audits the reviewer. It doesn't re-read the code. It checks that the review actually happened, with an approval, a completed template and green tests, and posts the checklist.
What it produced
By the company's account, over about five months the pipeline landed well over a million lines of code, merged 1,238 pull requests and needed only seven reverts. It consumed just over $135,000 in tokens, about $84 per issue. Every issue kept its own record of the plan, the failing reproduction test and each review, so the reason any change was merged can be reconstructed afterwards.
Their showpiece was adding support for IBM's Db2 database in under two days for a $4,172 token bill. They compare that with roughly nine months and $160,000 of engineering time for a similar feature in 2024. One caveat is theirs, and it is a big one: nobody reviewed the code during that first test, and they say they are still confirming it is correct. After the first week they switched on a mode that requires a human to approve every merge.
Where it struggled
- Bureaucracy. A one-word interface wording change went through five rounds of rework, and a reviewer once blocked a change over a typo in its description.
- Loops that don't end. An urgent one-line fix took eleven rounds over two days. In one review window, three tangled issues consumed about $631 of the $1,046 spent. They added a circuit breaker that hands over to a senior agent after repeated rounds, and say it is too early to measure its effect.
- A self-feeding backlog. Nearly half of all issues were filed by the agents themselves, and the agent that proposed new work had to be throttled because humans couldn't review its output.
- Instructions that rot. An audit of 25 instruction files found roughly 23% of about 100,000 words could be cut without changing any rule.
- Nobody learns. The authors admit the system suits experienced engineers and doesn't yet teach juniors much.
A one-person version
This part is our adaptation, not the company's advice. Most of the design translates to a solo developer working with an AI assistant:
- Ask for a written plan first, and read it before any code is written. This is the cheapest place to catch a mistake.
- Keep each change small enough to read in a single sitting. The pipeline capped changes at about 1,000 lines and split anything bigger.
- Don't let the assistant touch a test to make it pass. If a test must change, make that its own commit.
- Check that the new tests fail when the change is removed. A test that passes either way proves nothing.
- Set a limit before you start, such as three rounds, after which you step in yourself.
- Save the plan, the reproduction and the reason for each decision somewhere you can find later.
- Prune your instruction files from time to time. Rules pile up the way code does.
What it doesn't show
- The numbers are the company's own, and nobody outside has checked them.
- The Db2 result came from code that no human reviewed at the time.
- The system was built and supervised by two very experienced engineers, and a human approval step now sits in front of every merge. That experience is part of what makes it work.
- The cost figures cover tokens only. Waiting time for human reviewers, which they say dominates how long an issue takes, isn't priced.
What to watch
- Whether the circuit breaker actually cuts runaway review loops.
- Whether the team can put some classes of issue back into fully autonomous mode.
- Whether other teams adopt the model. It was running in four of the company's repositories at the time of writing.
Token costs are the other thread here. If you want the budgeting side, see Why Cloud AI Bills Are So Hard to Budget.
Source
For disclosure: we make iPhone apps, OffgridStem, OffgridScribe, OffgridVox and OffgridCam.
More from Offgrid Studio
How AI Is Reading Old Archives, and What It Finds
What Is Jev? The AI Model That Refuses to Write Text