You can review AI-written code without reading every line by moving your attention from the diff to the evidence: read and approve the tests, re-run the checks yourself and compare the exit codes with what the agent claimed, scan which files changed, and have a second model try to break the change. What is left for your eyes is a short list of places where a mistake is expensive.
This is the order I use on my own repos. None of it needs a special tool.
Why can't you just read the diff?
You can, for one agent. Past that, review is the limit. Simon Willison, writing about running coding agents in parallel, put it this way: "the natural bottleneck on all of this is how fast I can review the results." Agents produce a 600-line diff in ten minutes. Reading 600 lines with care takes most of an hour, and attention fades long before the end.
Step 1: read the tests, and read them first
A test is a short statement of what the code must do. Twenty lines of tests can pin down two hundred lines of implementation. If the tests say the right things and they pass, most of the diff is already accounted for.
So the first question in a review is whether the tests say what you meant. Check three things:
- Do they cover the task as you described it? Compare them with your prompt, line by line if it was short.
- Could they fail? A test that asserts a mock was called, or compares a value with itself, passes for any implementation.
- Did existing tests change?
git diff main...agent/fix-login -- tests/. A deleted assertion or a newskipin an old test is the first thing to question.
This works best when the tests are written and approved before the code exists, so the agent cannot shape them around what it built. That workflow, including how to detect edits to approved tests, is in Tests first with coding agents.
Step 2: check the claims against the record
Every agent run ends with a summary. "All tests pass." "Lint is clean." "I verified the build." A summary is a claim. The record is what was run and what it returned.
Agents are not usually lying. Summaries go wrong in ordinary ways: the agent ran one test file and described it as the suite, the last run was before its final edit, or the command failed and the output scrolled past. The exit code does not have these problems. Zero means the command succeeded, anything else means it did not.
So run the checks yourself, in a clean checkout of the agent's branch, and print the code:
git worktree add --detach ../app-verify agent/fix-login
cd ../app-verify
npm ci
npm test; echo "test exit=$?" # a clean checkout also catches files the agent never committed
npm run lint; echo "lint exit=$?"
npm run build; echo "build exit=$?"
If you ran the agent non-interactively, you already hold part of the record. claude -p exits with code 0 on success and non-zero when the run fails, and with --output-format stream-json the final result lists any permission denials, which tells you what the agent tried to do and was refused.
Then compare, one claim at a time:
| The agent said | What to look for in the record |
|---|---|
| All tests pass | The full test command, run after the last edit, exit 0 |
| I added tests for the new behaviour | New test files in git diff --stat, and they fail if you revert the fix |
| No other files changed | git diff --name-only main...agent/fix-login |
| The build is clean | The build command, exit 0, in a fresh checkout |
Any contradiction stops the merge. One false claim also lowers how far I trust the rest of that run, so I read more of it by hand.
Step 3: look at the shape of the change
Before any line of code, look at the file list:
git diff --stat main...agent/fix-login
You asked for a redirect fix in the auth module. Did package.json change? The lockfile? CI configuration? A migration? Files in a module the task never mentioned? Each surprise is either a good reason you can find in a minute or a problem. This takes thirty seconds and finds the largest category of unwanted change, which is the agent doing more than it was asked.
Step 4: have a second model try to break it
The agent that wrote the code is a poor judge of it. It carries the reasoning that produced the change and tends to confirm it. A reviewer that sees only the diff judges the result on its own terms. Anthropic's best practices for Claude Code recommend exactly this: have a reviewer look at the diff in a fresh context before you count the work as done.
I go one step further and use a model from a different vendor, on the theory that two models trained separately are less likely to share a blind spot. If Claude Code wrote the change, Codex attacks it. codex exec runs in a read-only sandbox by default, so give it write access to the checkout and nothing more (flags from the Codex non-interactive mode page):
cd ../app-verify
codex exec --sandbox workspace-write "Run git diff main...HEAD and read the change. Find inputs or states where it gives a wrong result. For each one, add a test that fails and run it. Change nothing except test files. Report only failures you reproduced."
For a plain second opinion with no prompt of your own, codex exec review --base main reviews the branch against main (codex-cli 0.159.0).
In the other direction, start a new Claude Code session in the worktree and run its bundled /code-review skill, or give it the same prompt.
Ask for a failing test. An adversary's prose review is one more thing to read. A failing test is a fact: run it, see it fail, and you have found a bug without reading the implementation.
One caution, also from Anthropic's guidance: a reviewer told to find gaps will usually report some even when the work is sound. Ignore style remarks and hypotheticals. Count only what it can demonstrate.
What still needs your eyes?
Checks tell you the code does what the tests say. They do not tell you whether it should exist in that form. I still read these by hand, every time:
- The tests. They are the specification now.
- Anything that touches money, authentication, permissions or deletion. Tests rarely cover every way these go wrong.
- Database migrations and data backfills. They run once, on real data.
- New dependencies. A name, a version and a reason. Check the package is the one you think it is.
- Changes to configuration, CI and build scripts. These decide what the checks check.
- Public interfaces. Function signatures, API responses and file formats that other code depends on.
- Concurrency and retries. Tests that pass once say little about races.
Where Vakr fits
Vakr is a Mac app I built to do these steps for every run. Before the code is written, the agent writes failing tests and you approve them; later changes to those tests are detected. On review, what the agent said is checked against the commands it ran and their exit codes, and a contradiction holds the merge. An agent from a different vendor gets the diff with one job, to write a test that fails. Vakr is in early access with a waitlist, and it does not remove the list above: the tests and the risky parts still need you.
Questions
Is it safe to merge AI-written code you have not fully read?
It depends on what stands in for the reading. If you approved the tests, ran them yourself with exit code 0 in a clean checkout, checked the file list, and read the high-risk parts, the unread remainder is covered by evidence. If all you have is the agent's summary, no.
Why not trust the agent when it says the tests pass?
Because the summary is written from the agent's memory of the session, and it can be out of date or cover only part of the suite. The exit code of the test command, run after the last edit, is the record. Checking takes one command.
Does a second AI reviewer catch real bugs?
Sometimes, and it also raises false alarms. That is why I ask for a failing test. A test that fails is a real finding whatever model wrote it, and a reviewer that cannot produce one has told you something too.
How long should reviewing one agent task take?
For a small, well-specified task: a few minutes on the tests, the time it takes to run the checks, half a minute on the file list, and whatever the risky parts need. If it takes much longer, the task was probably too big for one run.
How Vakr compares with other tools for running agents: One, Orca, T3 Code, Agentbox and Conductor.