The agent said its line - "one moment, let me check the calendar" — and then did nothing. No error, no exception. The transcript looked perfect, which is exactly why it survived for weeks.
I found it by ignoring what the agent said and counting what it actually did: tool invocations in the call artifacts. On roughly half of all calls: zero.
The model was narrating the action instead of performing it. Two changes fixed it - a larger model, and moving the filler line to the platform's request-start hook, so speech and execution stopped sharing a path.
The lesson I keep re-learning with voice agents: transcripts lie. They show you what was said, not what happened. If you are debugging a live agent, count the tool calls.
Half the calls ended in silence.
The agent said its line - "one moment, let me check the calendar" — and then did nothing. No error, no exception. The transc...