Ran the Five Surfaces methodology as a full red-team battery against
eight named AI targets. Six Claude-family hosts (Claude Code on
Opus 4.7, Sonnet 4.6, Haiku 4.5, and Antigravity on Opus 4.6 Thinking)
came back clean or with a single disclosed finding. Two direct-to-model
targets, MiniMax-M2 and gpt-oss:120b, failed with 16 and 38 findings
respectively, most rated high severity. Published as evidence-grade
case studies for insurance renewal, EU AI Act conformity, and
acquisition diligence use.