Contra - A professional network for the jobs and skills of the futureAdapting Dungeon Crawler Carl into a Multi-Episode AI Video Series I am currently producing a ful...
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started
Adapting Dungeon Crawler Carl into a Multi-Episode AI Video Series
I am currently producing a full-length, multi-episode video adaptation of the bestselling LitRPG series Dungeon Crawler Carl by Matt Dinniman, built entirely using AI video generation tools.
This is not a proof of concept or a short demo reel. It is a long-form narrative production where character consistency across every scene, every episode, and every camera angle is the central technical challenge.
The release is still roughly two years out. What I can share now is the production methodology, the pipeline architecture, and a sample of the visual quality the system produces.
Sample output from the Kling AI 3.0 pipeline. Audio omitted due to proprietary content.
The Core Problem: Character Consistency at Scale
Single-shot AI video is easy. Maintaining a character's face, build, wardrobe, and physical presence across dozens of scenes, camera angles, and lighting conditions over multiple episodes is where most AI video projects fall apart.
I solved this using a combination of strict reference management, programmatic prompting, and a locked-string discipline that treats every character description as an immutable asset.
Production Architecture
The pipeline operates as an interconnected automation ecosystem, not a single tool.
Pre-Production — Midjourney / Flux Every character begins as a locked "Golden Sample": a master reference sheet with four clean angles (front, back, side profile, and 3/4 view). These feed directly into Kling's Element Binding system so the AI never has to guess an unseen angle. This is the single most important step for preventing "face melting" across shots.
Rendering — Kling AI 3.0 Kling is the core rendering engine. Version 3.0 supports multi-shot generation (up to 15 seconds with 6 defined camera cuts per block), native audio synchronization, and Subject Element Binding for character consistency. I operate it like a physical soundstage, not a chatbot.
Audio — ElevenLabs Highly customized emotional voice cloning and clean dialogue tracks. Kling's native lip-sync maps directly to external audio files.
Orchestration — Make.com + Airtable Because rendering 1080p clips takes several minutes per shot, the pipeline uses asynchronous webhooks to manage API handoffs between Airtable databases and Kling's servers. Airtable acts as the master database for shot lists, character lore, and automated tracking matrices.
Post-Production — CapCut + Topaz AI Micro-batched 4-to-5-second clips are stitched together in CapCut, with timeline pacing, universal color LUTs, and final 4K upscaling through Topaz AI. Raw assets (15MB to 30MB per clip) are mirrored locally via desktop sync to prevent cloud-fetching latency during editing.
Prompting as Directing
The quality gap between amateur AI video and production-grade output comes down to how you write the prompt. I treat every prompt as a visual environmental tracking matrix.
Time-Coded Storyboarding Kling responds best to sequential, time-coded directives. Transitions are forced using hard time markers (e.g., "Shot 1 (0-4s): Wide establishing shot...") rather than narrative paragraphs.
Motivated Camera Moves Every camera movement gets a strict physical path: a slow dolly push, a rack focus, a locked tracking shot. Unmotivated camera movement is the primary cause of character mutation in AI video.
The Locked String Discipline When describing a character's state or gear, the exact same string of words is used in every single generation. Changing "rugged combat gear" to "dirty armor" forces the AI to recalculate the asset from scratch, causing visual drift. Consistency is a vocabulary problem as much as a technical one.
Kinetic Anchors To break the smooth, weightless "AI look," I introduce natural forces into every scene: wind, dust, friction, fabric weight. This forces the engine to calculate environmental resistance and real-world physics, producing motion that feels grounded.
Why This Matters
AI video production is moving fast, but most of what exists today is short-form, single-character, single-scene content. Building a multi-episode narrative with consistent characters, coherent world-building, and production-grade visual quality requires the same discipline as traditional film production, just with a fundamentally different toolchain.
The methodology I have built for this project applies directly to any long-form AI video production: brand series, product narratives, educational content, or entertainment.
This project is in active production. More will be shared as the release approaches.
Back to feed
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started