Voice Draft Made for ChatGPT by Sankalp TiwariVoice Draft Made for ChatGPT by Sankalp Tiwari

Voice Draft Made for ChatGPT

Sankalp Tiwari

Sankalp Tiwari

š“š¢š­š„šž : š“š”šž š…šžššš­š®š«šž šˆ'š šš®š¢š„š š­šØ š…š¢š± š•šØš¢šœšž
6 weeks. 31 users surveyed. 6 interviews. 1 clear behavioral insight. Here's the product solution I designed from research to PRD. šŸ‘‡
š“š”šž š©š«šØš›š„šžš¦ (š«šžšœššš©):
Students don't use voice on ChatGPT because voice auto-submits the moment you stop speaking. No review. No edit. No control. For a student mid-exam-prep, one wrong transcription and the whole query is gone. That one failure creates long-term avoidance. The fix isn't better speech recognition. The fix is restoring control.
šŸ‘ š¬šØš„š®š­š¢šØš§ šš¢š«šžšœš­š¢šØš§š¬ šˆ šœšØš§š¬š¢ššžš«šžš:
→ Direction 1: Restore Control — editable voice drafts before sending → Direction 2: Reduce Disruption — quieter, less intrusive capture → Direction 3: Structured Thinking — pause/resume for multi-part queries
I chose Direction 1.
It's the only approach that directly eliminates the primary behavioral barrier. Directions 2 and 3 address secondary friction — they can be layered in as the feature matures.
š“š”šž š¬šØš„š®š­š¢šØš§: š•šØš¢šœšž šƒš«šššŸš­ šŒšØššž
Voice input redesigned from instant-submit to draft-first.
The new interaction loop: Speak → see transcription in real time → draft held for review → edit or add more → explicitly send
Voice now behaves like typing: controllable, reviewable, and structured.
šŠšžš² šŸšžššš­š®š«šžš¬:
→ Editable voice drafts — transcription becomes editable text before sending → Real-time transcription — builds trust, reduces fear of errors mid-speech → Pause and resume — supports natural breaks without losing the draft → Mixed input — supplement voice with typing in the same draft → Explicit send — voice is never auto-submitted
š‡šØš° šˆ'š š¦šžššš¬š®š«šž š¬š®šœšœšžš¬š¬:
Primary metric: % of study-related queries initiated via voice Guardrail: no regression in overall message send rate or session duration
A/B plan: Control → current instant-send voice Treatment → draft-first voice Split → 50/50 on study-segment users Duration → minimum 2 weeks to account for learning curve
šŸ‘ š­š”š¢š§š š¬ š­š”š¢š¬ š©š«šØš£šžšœš­ š­ššš®š š”š­ š¦šž ššš›šØš®š­ ššŒ š­š”š¢š§š¤š¢š§š :
→ Low feature adoption is almost never a discoverability problem. Dig into the behavioral mismatch first. → Non-goals matter as much as goals. I explicitly called out that this feature doesn't try to improve STT accuracy or replace typing — knowing what you're not solving keeps scope clean. → Trade-offs are decisions, not gaps. Draft-first adds one extra step to the send flow. I accepted that trade-off because 77% of users value accuracy over speed — and this step is the entire value delivery.
This is my first end-to-end PM project — market landscape → user research → wireframes → full PRD with A/B plan, metrics, open questions, and trade-off documentation. NextLeap
Like this project

Posted Aug 30, 2026

Designed Voice Draft Mode to improve voice input control for ChatGPT users.