I built an AI support-ticket router on Replit, then tested it against the real model. Three things broke that no mock would have caught.
A silent failure. The first run sent 33 of 40 tickets to "needs a person". The newer Claude model rejects the temperature setting, and my code swallowed the error, so every call quietly failed. Now every failed call is logged.
The model thinks before it answers. With a tight token limit, some tickets got no answer at all and some draft replies stopped mid-sentence. Both have room now, and a cut-off draft is never saved.
One incident looked like four bugs. Four Safari checkout failures each got a different issue key, so the "3+ tickets, same problem" alert never fired. Sorting tickets in arrival order and asking for keys that name the problem rather than the customer turned them into one flagged incident.
If you're adding an LLM feature to your product, test it on real inputs early. The failures are rarely where you expect.
The temperature one, on the very first run. It hid the other two, because every call was failing before the model could think or pick a key. Once I fixed it, tickets started coming back with no answer, which was the token limit. The Safari keys only showed up after both were fixed.
It was. The worst part is that it looked like caution: a failed call fell back to "needs a person", which is what the app is meant to do when it's unsure. Now a failed call is logged as an error, not treated as low confidence.
Then: 1 month, 1 template.
Now: 1 week, 15 templates.
Same designer, different workflow.
I wrote up how I built no-code.supply: fifteen website templates, each with its own brand, in HTML and React, and some in Framer too. Claude Code did most of the typing. I did the directing.
Take a vampire, a grim reaper and a bat for a midnight ride through a haunted town. Hold to boost over broken track, brake before the big hills, and grab candy to refill your boost. Push your luck too far and your passengers end up splattered across the screen.