I stress-test an existing LLM, AI feature, or model-driven workflow to find the failure modes that matter — before your users do. You provide access to the system, app, or API (or representative outputs), the expected behavior, relevant constraints, and any known concerns. I perform adversarial test design, edge-case exploration, failure-mode discovery, consistency testing, instruction and constraint testing, evidence capture, and failure classification. You receive a structured adversarial test set, a findings report, a failure taxonomy, honest severity and impact context (no exaggerated claims), reproducible examples of each failure, and recommended regression cases so the failures stay fixed. One fixed-price project. I don't promise a specific number of vulnerabilities — I promise a rigorous, evidence-backed picture of where your system actually breaks.