Iy built a full wellness product vlog with just one product image — here's the AI stack that made it possible
No studio. No camera crew. No voiceover artist. Just a single image of a red light therapy mask and a workflow that turned it into a complete, polished vlog.
When a health and wellness client needed a daily ritual video showcasing their red light therapy mask, I had one asset to work with: the product photo. That's it. No lifestyle footage, no model, no set.
So I leaned fully into AI — not as a shortcut, but as a deliberate creative toolkit. The goal was to build something that felt human, warm, and authentic. A vlog that invited viewers into a morning routine, not a product ad that screamed "generated content."
Here's what the final video combined: cinematic visuals generated from the concept, a natural AI voiceover that felt like a real person narrating their ritual, and a background score that matched the calm, glowy energy of the brand.
🛠 Tools used:
Grok AI → Video generation
Google AI Studio → Voice & audio processing
ElevenLabs → Background music
ChatGPT + Nanobanana → Visuals & concept images
The biggest lesson? The creative brief matters more than the tools. Knowing the feeling you want the video to evoke before you open a single app is what ties every AI output together into something coherent.
For this one, the brief was simple: soft morning light, intentional self-care, a 5-minute ritual that feels like a luxury you deserve. Every tool choice, every prompt, every audio layer was in service of that single feeling.
The result? A vlog that tells a story, built from one product image and a clear creative direction.
I'm curious — if you were building a brand video with just one product image and zero footage, which part of this workflow would you start with: the visuals, the voice, or the music? And why?
Experimented a bit today with Krea and image generation for a case study I’m putting together around AI EarPods connected to OpenAI.
The focus has been on creating fashion-forward product imagery and art directing a world that feels specific to the identity, rather than just generating nice-looking AI images.
The trickiest part has been product consistency. Especially getting the EarPods to actually sit snug in the ear. If you’ve worked through this process, you probably know the struggle 😅
Simply telling AI to “make it fit more snug or in the ear” doesn’t always work. It loves to reinterpret the product every time.
Still experimenting, but getting closer. If anyone has found a good workflow for keeping products consistent across AI-generated shoots, I’d love to hear it!
Is AI finally good enough to shoot a product demo on its own, or human touch is still needed?
I generated a 20s product demo with Seedance 2.5 in a single take providing it with one product image reference and an example video transcript (posted in the comments).
Do you disclose...