Tests first with a coding agent means the agent writes failing tests for the task before any code, you read and approve those tests, and only then does it write the implementation, with the test files locked. When it finishes, two facts decide the result: the approved tests pass, and their hashes show nobody changed them.
It turns review from reading a long diff into reading a few short tests. Below is the whole workflow. The git, node and shasum commands were run while writing this (Node 22, git 2.50); the Claude Code flags are from its documentation for version 2.1.296.
Why write the tests before the code?
When an agent writes code and tests together, the tests describe what the code does. If the code has the bug, so do the tests, and both pass. Tests written before the code can only describe what you asked for, because there is nothing else to describe.
The second reason is the one specific to agents. An agent told "make the tests pass" has two ways to get there: fix the code or change the tests. Most of the time it fixes the code. Now and then it loosens an assertion, adds a skip, or rewrites the expected value to match what its code returns, and then reports success. If the tests were approved and frozen first, that path is closed, and you can prove it stayed closed.
Anthropic's best practices for Claude Code suggest a version of this with two sessions: have one Claude write tests, then another write code to pass them.
The workflow
The example task: add a slugify function to a small Node project. Each task runs on its own branch in its own worktree (see Git worktrees for AI coding agents).
1. Ask for tests only.
git worktree add -b agent/slugify ../app-slugify main
cd ../app-slugify
claude -p "Write tests in tests/slugify.test.js for a function slugify(text) exported from src/slugify.js. It lowercases, replaces each run of non-alphanumeric characters with one hyphen, and trims hyphens from both ends. Use node:test. Do not create src/slugify.js. Do not write any implementation." \
--permission-mode acceptEdits
2. Read the tests. This is the review, so take it seriously. A good result for this task looks like this:
const test = require('node:test');
const assert = require('node:assert/strict');
const { slugify } = require('../src/slugify');
test('lowercases and joins words with hyphens', () => {
assert.equal(slugify('Hello World'), 'hello-world');
});
test('drops punctuation and trims hyphens', () => {
assert.equal(slugify(' Ship it, now! '), 'ship-it-now');
});
test('returns an empty string for only punctuation', () => {
assert.equal(slugify('?!'), '');
});
Fifteen lines. I can hold all of it in my head, and I can see what is missing (nothing about accented characters, which I decide I do not need yet). Edit the tests yourself or ask for changes until they say what you mean.
3. Confirm they fail, for the right reason.
node --test; echo "exit=$?"
The exit code must be non-zero, and the failure must be the missing module or a wrong result. A test that passes before the code exists proves nothing. Check the reason too. While writing this I first ran node --test tests/, which fails on Node 22 because the runner does not take a directory that way. It was red, but for a reason that had nothing to do with the task.
4. Approve: commit, tag, and record the hashes somewhere the agent cannot write.
git add tests && git commit -m "tests: slugify (approved)"
git tag tests-approved-slugify
mkdir -p ~/approved
git ls-files -z tests package.json | xargs -0 shasum -a 256 > ~/approved/slugify.sha256
The manifest lives outside the worktree on purpose. Include the files that decide how tests run, such as package.json and the test runner's config. An agent can make a suite pass without touching a test file by changing the test command.
5. Lock the tests and ask for the code. How to lock them is the next section.
claude -p "Implement src/slugify.js so the tests in tests/slugify.test.js pass. Do not modify anything under tests/. Run node --test and stop when it passes. If you cannot make a test pass, stop and explain why." \
--permission-mode acceptEdits --allowedTools "Bash(node --test)"
6. Check both facts yourself.
node --test; echo "exit=$?"
shasum -a 256 -c ~/approved/slugify.sha256
git diff --quiet tests-approved-slugify -- tests && echo "tests untouched"
Exit 0 from the tests, OK on every line of the hash check, and no difference from the tag. If all three hold, the code does what the tests you approved say, and you have evidence for it.
How do you stop the agent editing the tests?
In three layers, and the last one is the one you rely on.
Tell it. "Do not modify anything under tests/" in the prompt works most of the time. That is not good enough on its own, which is why the other two exist.
Deny the edit. In Claude Code, add a deny rule to .claude/settings.json in the worktree once the tests are approved:
{
"permissions": {
"deny": ["Edit(tests/**)"]
}
}
Write the rule with Edit. The permissions documentation says file permissions are checked against Edit(path) and Read(path) rules only, so a path rule on Write is accepted and never consulted. Deny rules block in every permission mode. They cover Claude's file tools, file commands that Claude Code recognises in Bash such as sed and tee, and shell redirects. The same page is clear about the gap: they do not cover a subprocess that writes files indirectly, such as a Python or Node script.
If you want the agent to get an explanation when it tries, a PreToolUse hook on Edit|Write that exits with code 2 blocks the call and passes its message back to Claude. The hooks guide has a ready-made script for protected files.
Detect it. Because of that gap, and because other agents have different controls, prevention is never complete. Detection is. A SHA-256 hash of a file changes if a single byte changes, and the manifest sits where the agent has no reason to look. When I append one line to the test file, the check says so:
$ shasum -a 256 -c ~/approved/slugify.sha256
tests/slugify.test.js: FAILED
shasum: WARNING: 1 computed checksum did NOT match
$ echo $?
1
A failed check does not always mean bad intent. Sometimes a test was wrong and the agent was right to say so. That is a decision for you: read the proposed change with git diff tests-approved-slugify -- tests, and if you agree, approve again and record new hashes.
What tests-first does not give you
- It only covers what the tests say. If you approve weak tests, you get code that passes weak tests. Reading the tests well is the job.
- Some work has no cheap test. Layout, copy, performance under load, a migration on production data. Review those the slower way.
- Passing tests say nothing about what else changed. Look at
git diff --stat main...agent/slugifyfor files you did not expect. The rest of the review is in How to review AI-written code without reading every line. - It costs a round trip. Two runs and a pause for approval. For a one-line fix it is too much. For anything you would otherwise read for twenty minutes it pays for itself.
The pause is also what makes this fit unattended work. Approve tests for three tasks before you leave, and the implementation runs can go overnight; see How to run Claude Code overnight, safely.
Where Vakr fits
Vakr, the Mac app I make for people who run Claude Code and Codex, has this workflow built in. A task dispatched tests-first stops after the failing tests and puts them in the list of things waiting on you. You read a few short tests and approve. Only then does the agent write the code, and it cannot change the tests you approved: tampering is detected. Vakr is in early access with a waitlist, and tests first is part of the free tier. The steps above do the same thing by hand.
Questions
Can't the agent just change the tests to make them pass?
It can try. A deny rule on the test folder blocks the usual ways, and a hash manifest stored outside the repo detects every way. If the hashes no longer match what you approved, the run does not count as passing.
Does this work with Codex and other agents?
The workflow does, since it is two prompts, a commit and a hash check. The way you lock files differs by agent, which is one more reason to rely on the hash check for proof.
What if a test I approved turns out to be wrong?
Then the agent should stop and say so, and you decide. Read the proposed change to the test, approve it if it is right, and record new hashes. The rule is that tests change through your approval and never as a side effect of making them pass.
How Vakr compares with other tools for running agents: One, Orca, T3 Code, Agentbox and Conductor.