A text chatbot can take three seconds to reply and nobody minds. A voice agent can't. Leave a caller in silence for a beat too long and they say "hello?", start talking over it, or hang up.
That one fact shaped everything we built on Talk-Lee, an AI voice agent that answers business calls for healthcare, real estate and finance teams. Scheduling, support, lead qualification, around the clock.
The goal was a reply in under 500ms. You don't get there with a faster model. You get there by making sure nothing waits for anything else.
→ Speech to text streams while the caller is still talking
→ An LLM and intent layer keeps track of what they actually want
→ Text to speech streams back, so the agent starts talking before the whole answer is ready
→ An orchestrator decides in real time whether to answer, book or hand off to a person
Then it has to do something useful. It books into Calendly, logs the lead in HubSpot, and passes the hard calls to a human with the context already attached.
Where it landed. Under 500ms responses, 1,000+ concurrent calls, 30+ languages, GDPR and TCPA compliant.
Third project in a row with the same lesson. The model is the easy part. The plumbing around it decides whether anyone keeps using it.
Honest question for anyone running a sales or front desk line. Would you let an AI agent pick up your calls today? Booking only, booking plus qualifying leads, or nothing that talks to a customer yet?
This project focused on building and testing a practical AI security assessment lab for evaluating LLM defenses against prompt injection and jailbreak attacks.
I integrated Spikee by Reversec with a locally hosted cybersecurity model running through LM Studio, then added NVIDIA NeMo Guardrails to compare model behavior under three conditions: no guardrails, input filtering, and combined input/output protection.
The work included configuring the local model environment, building a custom FastAPI gateway, integrating NeMo Guardrails, troubleshooting model latency and timeout issues, creating a reusable Spikee target, and analyzing attack results using Spikee’s built-in reporting tools.
The project also explored different adversarial testing approaches, including prompt injection datasets, obfuscation, encoded attacks, Best-of-N testing, synthetic canary leakage tests, and structured benchmark comparisons.
The objective was to measure how much the guardrails reduced successful attacks while keeping the model, dataset, and testing conditions consistent.
This project demonstrates a hands-on approach to LLM red teaming, AI safety testing, prompt-injection assessment, and guardrail validation for organizations deploying generative AI systems.
𝐌𝐲 𝐫𝐨𝐥𝐞: Python Backend Engineer focused on AWS serverless architecture
𝐏𝐫𝐨𝐣𝐞𝐜𝐭 𝐝𝐞𝐬𝐜𝐫𝐢𝐩𝐭𝐢𝐨𝐧:
I built a serverless backend for an AI-driven trading platform handling real-time market data and analytics. I used AWS services like Lambda, API Gateway, RDS, and S3 to create a system that scales without manual infrastructure management. I designed pipelines for ingesting and processing live data to support trading signals and sentiment analysis. My focus was on keeping latency low, handling high concurrency, and ensuring the platform remained stable for global users.