AI Breaktesting: Find Where Your AI Actually Fails by Toby GoddenAI Breaktesting: Find Where Your AI Actually Fails by Toby Godden
AI Breaktesting: Find Where Your AI Actually FailsToby Godden
Cover image for AI Breaktesting: Find Where Your AI Actually Fails
I test AI systems against real human intelligence. Not synthetic benchmarks, not adversarial prompts designed to make a model say something dangerous. Real tasks, real complexity, real expectations.
Most AI evaluation happens inside the engineering bubble: standardised tests, leaderboard scores, cherry-picked demos. Breaktesting comes from outside that bubble. I talk to your system the way a human talks to a colleague, with the kind of messy, embodied, intuitive tasks that humans handle naturally. Then I notice where it breaks and document why it matters.
A hand-drawn cipher puzzle from 2005 defeated both Claude and ChatGPT. A wooden toy designed for toddlers broke frontier image generation. These aren't edge cases. They're the gap between what AI appears to understand and what it actually understands.

Who this is for

AI companies who want honest capability assessment from outside the engineering team. Your benchmarks say one thing; a breaktester with twenty years of cross-disciplinary creative practice says another.
Businesses adopting AI who need reality checks before trusting critical tasks to systems that hallucinate with confidence.
Education and policy teams who need evidence-based insight into what AI can and can't do, grounded in real human craft rather than synthetic benchmarks.
Creative industries looking for clarity on what AI actually threatens and what remains firmly, beautifully human.

How it works

You give me access to the models and a brief. I design targeted probes drawn from real creative and cognitive tasks: spatial reasoning, hand-drawn symbol reading, contextual memory, cultural knowledge, embodied logic. I document every failure pattern, score each assessment, and deliver plain-language findings your whole team can act on.
FAQs

Toby's other services
Starting at$130 /hr
Tags
AI Red Teaming
AI Evaluation
AI Research
AI Testing
Creative Technology
Service provided by
Toby Godden England, UK
2
Followers
AI Breaktesting: Find Where Your AI Actually FailsToby Godden
Starting at$130 /hr
Tags
AI Red Teaming
AI Evaluation
AI Research
AI Testing
Creative Technology
Cover image for AI Breaktesting: Find Where Your AI Actually Fails
I test AI systems against real human intelligence. Not synthetic benchmarks, not adversarial prompts designed to make a model say something dangerous. Real tasks, real complexity, real expectations.
Most AI evaluation happens inside the engineering bubble: standardised tests, leaderboard scores, cherry-picked demos. Breaktesting comes from outside that bubble. I talk to your system the way a human talks to a colleague, with the kind of messy, embodied, intuitive tasks that humans handle naturally. Then I notice where it breaks and document why it matters.
A hand-drawn cipher puzzle from 2005 defeated both Claude and ChatGPT. A wooden toy designed for toddlers broke frontier image generation. These aren't edge cases. They're the gap between what AI appears to understand and what it actually understands.

Who this is for

AI companies who want honest capability assessment from outside the engineering team. Your benchmarks say one thing; a breaktester with twenty years of cross-disciplinary creative practice says another.
Businesses adopting AI who need reality checks before trusting critical tasks to systems that hallucinate with confidence.
Education and policy teams who need evidence-based insight into what AI can and can't do, grounded in real human craft rather than synthetic benchmarks.
Creative industries looking for clarity on what AI actually threatens and what remains firmly, beautifully human.

How it works

You give me access to the models and a brief. I design targeted probes drawn from real creative and cognitive tasks: spatial reasoning, hand-drawn symbol reading, contextual memory, cultural knowledge, embodied logic. I document every failure pattern, score each assessment, and deliver plain-language findings your whole team can act on.
FAQs

Toby's other services
$130 /hr