Breaktesting: Human Intelligence vs AI Systems by Toby GoddenBreaktesting: Human Intelligence vs AI Systems by Toby Godden

Breaktesting: Human Intelligence vs AI Systems

Toby Godden

Toby Godden

Breaktesting is the practice of finding where AI systems fail against real human intelligence — not through standardised benchmarks, but through the kind of messy, embodied, intuitive tasks that humans handle naturally.
Red-teaming asks: can we make it say something dangerous?
Benchmarking asks: how does it score on this test?
Breaktesting asks: can it actually do the thing?
The answer, surprisingly often, is no.

Case Study 1: The Pirate Code

In 2005, I designed a hand-drawn cipher puzzle called Vibration 13: Le Code du Pirate. A custom alphabet of 13 symbols, each representing two letters distinguished by a dot. Three obscure geographical locations encoded in the cipher. A triangulation problem leading to a tiny Indonesian island called Natuna Besar.
I deployed it as a Flash interactive. Thousands played. 274 people solved it.
In February 2026, I tested both Claude (Anthropic's Opus 4.6) and ChatGPT against the same puzzle.
Neither solved it.
ChatGPT spent over thirty minutes cropping and zooming the image, burning significant computational resources, and produced confident but wrong symbol readings. It hallucinated answers with full conviction. Score: 3/10.
Claude admitted early that it couldn't reliably read the hand-drawn symbols. It understood the cipher logic when I explained it, verified the geography when I provided decoded place names, but could not perform the core task: reading symbols from a drawn page. It was honestly useless rather than confidently wrong. Score: 2/10, Claude's own self-assessment.
A puzzle designed twenty years ago, solvable by hundreds of humans with an atlas and a pencil, defeated both frontier AI systems comprehensively.

Case Study 2: The Puzzle Pebble

My father, Reinold "Rhino" Godden, was a toymaker in Ystradgynlais. In 1996 he designed the Puzzle Pebble: a smooth wooden disc cut into six interlocking jigsaw pieces, painted in bright colours.
One Pebble is a tactile learning toy a child can solve. Add more Pebbles and the difficulty scales exponentially: each pebble's six pieces are uniquely cut with different interlocking profiles, so they're not interchangeable, but pieces from different pebbles almost fit together, creating productive deception. You have to figure out which pieces even belong together before you can solve any of them.
I asked multiple AI image generators to draw the Puzzle Pebble, both assembled and disassembled.
They can draw a whole disc. But the moment you ask for the six separated pieces, they produce what I call broken biscuits: vaguely curved fragments with no consistent geometry, no way to reassemble them into a circle. The spatial logic that a six-year-old grasps intuitively is beyond current generative AI.
A toy designed for toddlers breaks frontier image generation.

Case Study 3: The Small Things

Breaktesting isn't always dramatic. Sometimes it's noticing, mid-conversation, that an AI system can't tell you what time it is. Or that it confidently states a fact about your own work that it's fabricated. Or that it handles a Welsh place name fine in one message and mangles it in the next.
The value is in paying attention. Using the system naturally, noticing the wobble, then designing a targeted probe around it. Most users shrug and move on. A breaktester documents and drills down.

Why This Matters

The gap between what AI appears to understand and what it actually understands is being papered over by confident language and slick demos. Breaktesting cuts through that.
For AI companies: Honest capability assessment from outside the engineering bubble. Your benchmarks say one thing; a hand-drawn puzzle from 2005 says another.
For businesses adopting AI: Reality checks before you trust critical tasks to systems that can't read a dotted symbol or draw six arcs.
For education and policy: Evidence-based insight into what AI can and can't do, tested against real human craft rather than synthetic benchmarks.
For creative industries: Clarity on what AI actually threatens and what remains firmly, beautifully human. More than the hype suggests.
Like this project

Posted Jul 28, 2026

Finding where AI systems fail against real human intelligence. A cipher puzzle from 2005, a toymaker's wooden disc, and the small failures that benchmarks miss