You spend an hour getting a feature working with an AI. The code looks fine in the pull request. Variable names make sense. Error handling is there. Types check out. You approve it. It goes live. Then it crashes because a webhook sent a null user_id something that "should never happen."
Here is the problem. AI code reads well because the examples it learned from are all clean and perfect. It falls apart in real situations. The model doesn't know your system.
Two things it misses:
It only knows the success case. The model writes code for when everything works. Every tutorial shows the success case. It doesn't write code for when the request gets cut off halfway. That never appears in examples. It only appears in your logs at 2 AM.
Assumptions you can't see. The code assumes the payment API returns "status: ok". It assumes the cache has data. It assumes the list isn't empty. These look like normal defaults. But in your system, the payment API returns "data.status: succeeded" on success. On timeout it throws a 500 with HTML garbage.
How to catch this:
Test the empty case first. Send nothing. Send an empty array. Leave out the header. Use an expired token. The happy path works in the demo. The empty path breaks in production.
Then read the code and ask: what does this assume is true? Not what it does what it assumes. The API format. The data shape. The timing. The permissions. The timezone. Check each one against your actual system. Not the one the model imagined.
The code isn't bad. It was written for a perfect world. Your world isn't perfect. The review that matters happens in your codebase with your data. Not in the chat where the code was written.
This is the kind of AI workflow that gets me excited as an independent developer.
Faster ideation, faster prototyping, and more room to focus on turning ideas into real products. 🚀