LLM Security Assessment Lab for Prompt Injection and GuardrailsLLM Security Assessment Lab for Prompt Injection and Guardrails
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started
This project focused on building and testing a practical AI security assessment lab for evaluating LLM defenses against prompt injection and jailbreak attacks.
I integrated Spikee by Reversec with a locally hosted cybersecurity model running through LM Studio, then added NVIDIA NeMo Guardrails to compare model behavior under three conditions: no guardrails, input filtering, and combined input/output protection.
The work included configuring the local model environment, building a custom FastAPI gateway, integrating NeMo Guardrails, troubleshooting model latency and timeout issues, creating a reusable Spikee target, and analyzing attack results using Spikee’s built-in reporting tools.
The project also explored different adversarial testing approaches, including prompt injection datasets, obfuscation, encoded attacks, Best-of-N testing, synthetic canary leakage tests, and structured benchmark comparisons.
The objective was to measure how much the guardrails reduced successful attacks while keeping the model, dataset, and testing conditions consistent.
Tools used: Spikee, NVIDIA NeMo Guardrails, LM Studio, Python, FastAPI, PowerShell, Parrot OS, local LLMs, JSONL datasets, and custom security testing scripts.
This project demonstrates a hands-on approach to LLM red teaming, AI safety testing, prompt-injection assessment, and guardrail validation for organizations deploying generative AI systems.
Moch Virgiawan's avatar
love that you actually red-team it instead of shipping and praying 😅
Stephen's avatar
Even with the guardrails enforced, some attacks still bypassed it
Moiz Ahmed's avatar
pretty cool tbh. testing the same attacks with different guardrail setups makes the results way more interesting.
Stephen's avatar
Yes, it does. The focus next is experimenting with different attacks and seeing how the guardrails work
Moiz Ahmed's avatar
Yeah, that sounds like a good direction. Would be interesting which attacks still get through once you start pushing the guardrails harder.
Kapil's avatar
Holding the model and attack set constant makes this comparison useful. Since some attacks still got through, did you also measure false positives on legitimate requests and latency for each guardrail setup? Those two numbers often decide whether a defense survives production traffic.
Stephen's avatar
Thanks for the insight. Will do that
Back to feed
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started