MCP-Powered Voice Agent for Database and Web Data RetrievalMCP-Powered Voice Agent for Database and Web Data Retrieval
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started
MCP-Powered Voice Agent: AI Speech-to-Text & Data Orchestration.
I recently engineered a Model Context Protocol (MCP) driven voice agent for a client who needed a seamless, hands-free way for their users to query internal databases and retrieve real-time internet data. The client faced a significant challenge: their users were experiencing friction when manually searching through complex internal records, and existing text-based chatbots were too slow and cumbersome for their fast-paced, on-the-go operational environment. They needed a highly responsive, voice-native AI solution capable of intelligently routing queries between proprietary data and external web searches without requiring manual intervention.
To solve this, I designed a real-time voice orchestration pipeline utilizing LiveKit for seamless audio management. The workflow begins when a user speaks into the system, which immediately processes the audio through AssemblyAI for highly accurate Speech-to-Text transcription. Once transcribed, the text is routed to the Qwen3 Large Language Model (served locally via Ollama), which acts as the central reasoning engine to interpret the text. By leveraging the Model Context Protocol, the agent can dynamically discover and evaluate available tools based on the exact intent of the user's spoken request.
The core implementation relies on intelligent, autonomous tool invocation. If the Qwen3 LLM determines the user's query is related to internal records, it securely invokes the right tool to query the client's Supabase database via dedicated MCP connections. If the required information is not found internally, or if the user asks a broader question, the system automatically falls back to utilizing Firecrawl to perform a live web search. After fetching the necessary data from either source, the LLM generates a cohesive text response, which is instantly converted back into natural audio via a Text-to-Speech module and delivered as speech output to the user.
The client was highly satisfied with the final deployment. The MCP architecture provided an incredibly robust and scalable framework, allowing them to easily and securely expose new database tables to the AI in the future without rewriting the core voice logic. By effectively eliminating the manual search bottleneck and delivering near-instantaneous, accurate voice responses, the project successfully modernized their data retrieval process and dramatically improved overall workflow efficiency.
Post image
Back to feed
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started