AI-Based Synchronized Video Dubbing
Owais Ansari · · International Conference on ICT for Sustainable Development
Abstract
A multilingual video dubbing system that chains automatic speech recognition, neural machine translation and text-to-speech synthesis while preserving audio–video timestamp alignment. The system supports ten Indian languages and runs on a hybrid backend: a Node.js API gateway in front of Flask services that host the AI inference workloads.
What the paper does
Owais Ansari published research on a multilingual video dubbing system at ICT4SD 2025, with proceedings published by Springer. The system integrates OpenAI Whisper for automatic speech recognition, a neural machine translation pipeline, and text-to-speech synthesis, and supports ten Indian languages — a setting where off-the-shelf dubbing tooling is thin.
Dubbing is not translation with a speaker attached. Each stage degrades the timing of the one before it: recognition produces segments whose boundaries do not align to sentences, translation changes utterance length unpredictably, and synthesis imposes its own prosody and rate. Chain them naively and the audio drifts away from the speaker’s mouth within a minute.
The synchronisation problem
The contribution is maintaining accurate audio–video timestamp synchronisation across that chain. Whisper emits timestamped segments, which gives the pipeline an anchor, but translation between languages routinely changes duration — and the ten supported Indian languages differ substantially from one another in how compactly they express the same content. The system has to reconcile translated audio against the original timeline rather than assume they match.
Why the backend is split
Architecture follows the workload. A Node.js API gateway handles request routing, orchestration and client-facing concerns, while Flask services host the AI inference stages. That split lets the I/O-bound coordination layer and the compute-bound model layer scale on their own terms — the gateway stays responsive under concurrent requests while inference runs where the Python ML ecosystem already lives, with no bridge between runtimes.
Methods
- OpenAI Whisper
- Automatic speech recognition
- Neural machine translation
- Text-to-speech synthesis
- Timestamp synchronisation