Engineering Practices · AI & Testing
AI Coding isn't risky. Missing Tests Are.
AI made shipping faster. Approval tests make it safe: high coverage on legacy code for little effort, and a feedback loop your AI agent can use.
The Speed Is Real
I've been chasing development speed for 15 years. First with ReSharper, then with TDD, and now with Claude Code.What changed is the size of the jump. Features that took days now take hours. Refactors nobody dared to touch get done in an afternoon.
But there is one question that rarely comes up when people show off how fast they are shipping:
How do you know it still works?
Faster Code, Same Eyes
AI writes code faster than any of us can read it. So we don't read all of it. We skim the diff, run the happy path once, and ship.That's fine for a weekend project. On a product with paying users, every change you can't verify is a bet, and you are not betting your own money.
AI without automated tests is gambling, and the chips belong to your users.
None of this is new. The book Accelerate, by Nicole Forsgren, Jez Humble and Gene Kim, showed it years ago: the teams that deliver the fastest are also the most stable, and they get there through practices like test automation and continuous delivery.
Delivering quickly without the right practices generally ends in tears. AI raised the speed for everyone, but the practices still have to come from you.
With a good test suite, the same speed looks very different. The AI makes a change, the tests run in seconds, and a red test tells you what broke. The AI can even read the failure and fix it by itself.
That feedback loop is what lets you stay fast for longer than one sprint.
The Luxury I Rarely Found
As a TDD practitioner, delivering was always easy on codebases where my team had good test coverage. Change something, run the tests, deploy. No drama.But when joining new environments, that luxury was rarely there. Years of business rules, no tests, and the people who wrote them long gone.
Now put AI on top of that: an assistant that confidently changes 30 files in a codebase where nothing checks the result.
You Have two ways to make changes now:
Change it surgically
Touch as little as possible, test by hand, and pray. It works until the day it doesn't. It also throws away the speed AI gave you, because a human has to check every change manually.
Or add test coverage first
The right answer, but it sounds expensive. Unit testing code that was never designed to be tested can take months, and you often have to refactor it without a safety net just to make it testable.
So how do you get a high amount of coverage for very little effort?
The Hadouken of Tests
There is a secret technique for this: approval tests, also known as golden master or snapshot tests.
With these tests you only care about what goes in and out of the system. How it works internally doesn't matter.
- Give the software its inputs: an HTTP request, a message, a file.
- Spy on its outputs: the response, the rows it tries to save, the messages it tries to publish.
- Take a snapshot of how it behaved, and approve it.
From then on, every run compares the new output against the approved snapshot. Any difference fails the test and shows you the diff.
Take an endpoint that places an order. You send an order in, then capture the response, the order it tries to save, and the message it tries to publish.
The first run records that behaviour. You read it, approve it, and that is your safety net.
The snapshot records what the system does today, bugs included. That is what you want before letting anyone, human or AI, change it.
Don't trust it blindly, though. Invert an if, comment out a line, and check that the test goes red. If it stays green, you found a gap.
And here is the nice part: writing these scenarios is a perfect job for the AI. It generates the inputs, you review the snapshots.
I went into detail on this technique here.
Why We Mock the Database and the Queues
You could run these tests against a real database and a real queue. I don't, and it is a trade-off."There are no solutions, only trade-offs." Thomas Sowell
When a test calls real dependencies it gets slow, it needs data to be set up first, and it fails because of the network. Slow and flaky tests stop being run, and then you are back to gambling.
So we mock at the very last level: the class that talks to SQL, the client that publishes to the queue, the HTTP call to another team's API. Everything else is the real code, running in-process.
That gives us:
- Speed: nothing leaves memory, so the whole suite runs in seconds.
- Parallel runs: there is no shared database, so tests don't step on each other.
- No false alarms: a red test means the behaviour changed, never that a dependency was down.
- More to spy on: the fakes record what the code tried to save and publish, and that goes into the snapshot.
The price is that the thin layer talking to the real database stays out of this coverage. Cover it separately with a handful of integration tests.
Speed matters even more now. An AI agent runs the tests after every change. If your suite takes 20 minutes, your fast AI waits 20 minutes. If it takes 10 seconds, it gets feedback at every step.
A more in-depth guide on this test scope is here.
To Wrap It Up
AI gives you the speed. The tests decide whether that speed is delivery or gambling.If your codebase has no coverage, don't start with the feature. Start with a handful of approval tests around the area you are about to change, then let the AI loose.
Happy coding.