Client: Aphra
Role: AI Developer
Team: Founder, Muhammad and two other developers
Duration: Approximately six to twelve months
Product: Personal AI assistant with live voice and an interactive avatar
Website:https://aphra.me
Aphra was an early personal AI product where people could talk to an assistant that remembered their context and helped with daily work over time.
I worked directly with Aphra's founder as the AI Developer in a three-developer team. My responsibility was the AI interaction layer. I built the live voice pipeline, agent orchestration, persistent memory and avatar integration in Python. The other developers owned the frontend, general backend, OAuth connections and cloud infrastructure.
The first major version, including the live talking avatar, was rolled out in about nine weeks. I continued building and optimizing the AI capabilities for roughly six to twelve months.
Building real-time conversation before real-time model APIs
I built this before we had a real-time LLM API for the complete conversation path. The voice experience had to be assembled from separate systems.
I built a WebSocket-based pipeline that connected:
Live user audio
Speech-to-text
LLM and agent orchestration
Persistent user memory
Text-to-speech
Avatar animation and lip synchronization
Each stage added delay. I worked on the orchestration and timing so the response could move through the complete path in under a second and remain usable as a live conversation.
Lip synchronization was a separate technical problem. I integrated the avatar and synchronized its facial movement with the generated speech while keeping the interaction responsive.
Turning an LLM into a persistent personal agent
Aphra needed to remember a person across conversations and continue work that lasted longer than one response.
I built the memory and orchestration layer using Python, LangChain and LangGraph. It retained relevant user context and kept long-running agent work from losing its place. This allowed the product to support more than question-and-answer chat.
Users could talk through new ideas, keep notes, review meetings, create to-dos, draft emails and discuss their daily routines. The assistant carried relevant context into later interactions.
That persistence was the hardest part of the project. Agent frameworks were still early, so I had to design much of the memory and operating logic inside the product. It needed to retain context and stay stable while an interaction or task continued over time.
Keeping the AI layer stable under real usage
The AI integration eventually handled voice calls, chat, documents and long-running agent work across the product. The platform reached a one-time recorded user peak of 16,899.
The agent layer continued to retain context during long-running work without requiring a rewrite. The product reported 99.8% uptime, and the conversational path remained responsive enough to preserve the live-avatar experience.
The AI was the interface people spoke to and used for ongoing work. A lost memory or broken session was immediately visible to the user.
Results
A one-time recorded user peak of 16,899.
The first live-avatar feature set rolled out in approximately nine weeks.
Sub-second conversational response through speech recognition, agent reasoning, voice generation and avatar output.
99.8% reported uptime for the product. The original measurement period is not available.
Long-running agent work retained user context and stayed stable during the reported user peak.
Live voice and avatar interaction built before a single real-time model API covered the complete pipeline.
Aphra combined live avatars, voice and continuing personal context in one assistant experience. This case study covers the historical AI-avatar layer I built.
The interaction path joined live audio, speech recognition, agent memory, model orchestration, voice generation and avatar lip synchronization while staying responsive.
Continued AI development and optimization for approximately six to twelve months.
My exact scope
I owned the Python AI layer: WebSockets, speech-to-text, LLM orchestration, text-to-speech, agent memory, avatar integration and lip synchronization.
I did not own the product frontend, general backend, OAuth integrations, GCP infrastructure or infrastructure cost management. Those areas were handled by the other developers on the team.
A one-time recorded user peak of 16,899, a nine-week first live-avatar rollout, sub-second conversational response and 99.8% reported uptime.
Related service
I help teams build production AI agents with persistent context, live voice interaction and long-running workflows.
Built Aphra's live voice, agent memory and avatar AI in Python, supporting long-running assistant workflows and a one-time recorded peak of 16,899 users.