‘Jai Hanuman’ for STAGE: Inside India’s First AI Vertical Drama

 

ai vertical drama jai hanuman

Artificial Intelligence is rapidly transforming the way stories are created, but Jai Hanuman demonstrates that successful AI filmmaking is about far more than generating impressive visuals. In this exclusive interview, Manik Prajapati shares how his team developed one of India’s first AI vertical drama series for STAGE, available in multiple regional languages.

He explains the layered production pipeline behind the project—from story development and mythology research to character design, AI-generated visuals, multilingual voice creation, and post-production. The conversation also explores why Hanuman was chosen as the central character, how the vertical format was designed specifically for mobile audiences, and why cultural authenticity remained a non-negotiable priority throughout the production process.

Beyond discussing technology, Manik offers valuable insights into the future of AI-powered filmmaking. He candidly talks about the challenges of maintaining character consistency, preserving emotional depth, building custom AI workflows, and balancing automation with human creative judgment. He emphasizes that great AI productions are built on disciplined production pipelines rather than prompts alone, making this interview an insightful read for filmmakers, animators, VFX professionals, AI artists, students, and anyone curious about the next evolution of storytelling.

From Editor to AI

1. Welcome Manik Prajapati. Let us know about your professional journey.

I started as an editor. Cutting promos, trailers, the stuff that gets an audience to click play in the first three seconds. Over time the job kept pulling me toward the tools themselves rather than just using them. If a workflow was slow, I’d end up building something to fix it instead of waiting for someone else to. That habit is basically my whole career arc now — editor to tool-builder to what I’d call an AI infrastructure person.

At STAGE, that meant going from cutting promos to designing the actual production pipeline that made a series like Jai Hanuman possible in the first place.

Birth of Jai Hanuman

2. How did you get the idea of vertical ‘Jai Hanuman’?

STAGE’s whole audience lives on phones in Tier 2 and Tier 3 India, and mythology has always worked there because people already know the story beats. Hanuman specifically is a character with huge emotional range — devotion, rage, humor, grief — which made him a good stress test for whether AI tools could actually hold a character’s identity across an entire arc, not just one pretty shot.

The vertical format wasn’t an afterthought either. If the audience is watching on a five-inch screen during a commute, the whole grammar of the shot has to be rebuilt for that frame, not just cropped from a widescreen version.

Why regional languages?

3. Why is it published in several Indian languages?

Because “Indian audience” isn’t one audience. A Bhojpuri viewer and a Gujarati viewer don’t just speak differently, they respond to a scene differently — pacing, humor, how devotion is expressed. We didn’t want dubbed-over Hindi with a regional accent slapped on top, we wanted the dialogue actually written in that dialect’s grammar.

So the series exists in Hindi, Rajasthani and Haryanvi, Bhojpuri, Marathi, and Gujarati, each with its own pass rather than one master track translated five times.

Inside the AI Pipeline

4. Maintaining your actual production pipeline, please share what your production pipeline looks like.

It runs in layers, and each layer only feeds the one above it. Story and canon sit at the bottom — we kept Valmiki’s Kishkindha Kanda as the primary source with Puranic material as secondary, so nothing drifts from what the audience already trusts. Above that is pre-production: character design, voice, world rules. Then direction, which is shot breakdown and continuity. Then generation itself, where the actual frames and audio get produced.

Post-production sits on top, doing assembly, color, and final QC. The generation stage doesn’t get to make canon decisions, and the assembly stage doesn’t get to invent shots that weren’t planned. That discipline is what keeps a twenty-plus episode arc from turning into a different show by episode ten.

Making AI feel human

5. Being an AI script, how do you keep emotions authentic?

We built around the Navarasa system, the classical nine emotions from Indian aesthetic theory, instead of trying to reinvent an emotion framework from scratch. Every scene gets tagged to a dominant rasa before it’s generated, and that tag drives everything downstream — voice delivery, music mood, even how a shot is framed.

It’s not a Western six-emotion model bolted onto Indian content. It’s built from the tradition the content already belongs to, which honestly made the whole authenticity problem easier to solve than I expected going in.

AI tools / softwares used in ‘Jai Hanuman’

6. Let us know about the various AI tools used for storyboarding, animation, voice, etc.

Tools and their respective work are as follows.

  • ElevenLabs v3 handles voice, with emotion tagging tuned per character and per language since Haryanvi delivery and Bhojpuri delivery don’t hit the same beats.
  • Suno generates the score, mixing devotional, epic, and folk elements depending on the scene.
  • Seedance 2.0 is the main engine for the actual animation and cinematic shots.
  • Kling and Runway used for specific shot types where they outperform.

The toughest challenge

7. What was the hardest technical problem?

Character consistency across languages and across the Vanara characters specifically. Getting Hanuman to look like Hanuman in shot forty the way he looked in shot four is one problem. Getting an entire army of Vanara characters to read as the same species, with recognizable anatomy, rather than forty different monkey designs, is a much harder one. We ended up building a character-locking system with strict style anchors just to keep that consistent across thousands of generations.

Where AI fails?

8. Which AI limitation frustrates you most?

Emotional subtlety at the small scale. The tools are genuinely good at big, obvious emotion — grief, rage, joy read fine. What they still struggle with is the quieter register, a character who’s devoted but also tired, or angry but trying not to show it. That’s exactly the register a lot of Indian storytelling lives in, so it’s not a minor gap for this kind of content.

AI for Independent Cinema

9. Can AI make indie filmmaking easier?

Yes, but only for people willing to build real infrastructure around the tools rather than just prompting and hoping. Anyone can generate a pretty shot. Very few people are building the character-locking systems, the continuity checks, the dialect-accurate scripts that turn a pile of generated clips into an actual series someone will watch to the end.

The tools lower the cost of a single image. They don’t lower the cost of holding an entire filmmaking and that part still takes real production discipline.

Does AI save money?

10. Can AI actually reduce production cost and time?

For certain categories, yes, meaningfully. Music, voice, and a large share of the visual generation move much faster than a traditional animation pipeline. But there’s a real ceiling — canon accuracy, dialect grammar, character consistency, and quality control all still need trained human judgment on top.

What we’ve built isn’t “AI replaces the crew,” it’s AI collapsing the parts of the pipeline that were pure execution, freeing the humans to spend their time on the parts that actually need taste.

Human touch matters

11. What percentage of the final visuals are directly usable from AI versus manually refined?

This varies episode to episode, but a fair working range for us has been roughly 60–70% usable close to as-generated, with the rest needing manual refinement for continuity, color matching, or fixing anatomy errors on background characters. 

How many prompts?

12. Approximately how many prompts were generated for one finished episode?

It really depends on the scale and length of the episode, there’s no fixed number I can quote across the board. A shorter, simpler episode might come in under a hundred prompts total. A longer one with heavy action or multiple new characters can go well beyond that, sometimes by a wide margin, once you count iterations and rejected generations. It’s not something we’ve standardized because the episodes themselves aren’t standardized in scope.

Training custom AI

13. Did you create custom LoRAs or fine-tuned models?

For a few specific parts, yes, we did train custom LoRAs. It wasn’t the default approach everywhere, base model quality moves fast enough that a fine-tuned model can go stale within months, but for the pieces that needed a very particular locked look, we went ahead and trained. That work involved a lot of people’s effort behind the scenes, and honestly it wouldn’t have happened without our CEO pushing the vision for what this series needed to look like.

Character consistency for the rest was handled through disciplined reference anchoring and locked style parameters, so we weren’t stuck maintaining a model every time a better base model shipped.

Most underrated AI tool

14. Which underrated AI tool deserves more attention?

GPT’s image generation, honestly.

Most of the attention in this space goes to the dedicated video generators, but GPT’s image model quietly does a lot of heavy lifting for us on character design and reference frames, the kind of consistent, controllable output that everything downstream depends on. It doesn’t get talked about the way the flashy video tools do, but a shaky reference image ruins every generation that follows it, so this is one of those unglamorous tools that ends up mattering the most.

Managing 3D assets

15. How do you organize thousands of generated assets?

Honestly, by the time we’re generating, most of the episode is already mapped out in my head. The shots, the sequence, which scene needs which rasa, I’ve usually run through the whole episode mentally before a single prompt goes out. So organizing at that point becomes less about sorting through chaos after the fact and more about commanding the system to execute what’s already planned.

Everything still gets tagged at creation time, character, scene, language variant, take number, but that tagging is following a structure I already know, not discovering one.

Advice for future AI artists and creators

16. What advice would you give animators and filmmakers?

Don’t treat prompting as the skill. The prompt is the easy part. The actual skill is building the systems around the prompt, the consistency checks, the continuity rules, the quality gates, that turn a hundred lucky generations into a series someone can watch start to finish. Learn the craft of pipelines, not just the craft of prompts.

Building AI-Native OTT series

17. What advice would you give an OTT platform considering an AI-native series?

Start narrow before you go wide. Don’t greenlight a fifty-episode AI-native series as your first attempt, prove the pipeline on a small arc first, measure quality honestly, and only then scale it up. Build the layered architecture from day one too, story and canon at the base, then pre-production, then direction, then generation, then post, each layer feeding only the one above it. If you skip that structure to move faster early on, you pay for it later when a twenty-episode arc starts drifting from its own canon.

The bigger point for Indian platforms specifically: don’t treat cultural authenticity as a nice-to-have layered on top at the end. Dialect grammar, regional pacing, festival and mythological accuracy, these need to be part of your actual quality checks from the first episode, not a patch after audience complaints show up. An AI-native series that nails the visuals but gets the Bhojpuri or Haryanvi delivery wrong will lose the exact audience it was built for. And budget for the humans.

The tools cut the cost of generating a shot, they don’t cut the need for people with real production judgment sitting on top of the pipeline, checking continuity, catching what the model got wrong, and deciding what’s actually good enough to ship.