AI Agent Security Testing and Real-World Failure BenchmarkingAI Agent Security Testing and Real-World Failure Benchmarking
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started
I test AI agents for the failures that demos hide.
Over the past few months I've built Sentinel, a benchmark that puts MCP security guards through 6 classes of attacks in a sandbox, and I evaluate frontier coding agents on real bugs every week.
The biggest thing I've learned: an agent that looks great in a demo can still take unsafe actions, make up an answer, or quietly misread an ambiguous instruction once it meets real users.
If you're shipping an AI feature and want someone to break it before your users do, send me a message. I'll point out 3 issues I'd test first, free.
Back to feed
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started