Pixazo’s AI VFX Studio is an AI video translator that runs end to end: give it a video you already have and it will translate video dialogue into another language, clone the speaker’s own voice from that same clip, speak the translated line, and re-time the mouth to the new audio. It is free to start — a 7-day trial with no card and 100 bonus credits. In a measured English-to-Spanish run on 29 July 2026 a five-second, single-speaker clip took about fifteen minutes, and both clips are on this page with sound.
What does Pixazo’s AI video translator actually do to a video?
It keeps the picture and the speaker’s own voice, replaces the language, and re-times the mouth to the new audio — the two clips below are the genuine output of one run on 29 July 2026, and because the evidence is audio you have to press play.
Play the English clip first, then the Spanish one — one at a time. The audio is the entire claim.
Same face, same performance, different language.
EN · Source“Welcome to our community. We’re so glad you’re here.” — the English source clip, generated on Pixazo, so the voice in it is ours to clone.
ES · Dubbed“Bienvenidos a nuestra comunidad — nos alegra mucho que estés aquí.” — the Spanish dub from the same run, spoken in a voice cloned from the source clip and re-timed to the picture.
LanguageEnglish → Spanish
VoiceCloned from the clip’s own speech, no separate sample supplied
MouthRe-timed with sync_mode “remap”
What the agent printed before it generated any audio
pixazo_transcribe →
“Detected English, 5.68 seconds, 1 speaker”
transcript →
“Welcome to our community. We’re so glad you’re here.”
translation (written by the agent) →
“Bienvenidos a nuestra comunidad — nos alegra mucho que estés aquí.”
The clip file is about five seconds long; the transcriber reported 5.68 seconds of speech. We print both exactly as they came back and do not reconcile them.
Run log — Pixazo AI VFX Studio, 29 July 2026. One clip, one speaker, English to Spanish.
Stage
What the agent called
What came back
Measured time
01 Transcribe
pixazo_transcribe
“Detected English, 5.68 seconds, 1 speaker”, plus the timed transcript “Welcome to our community. We’re so glad you’re here.”
13 s
02 Translate
the agent itself — no separate translation model
“Bienvenidos a nuestra comunidad — nos alegra mucho que estés aquí.”, phrased to fit the spoken duration it had measured
not timed separately
03a Pull the source audio
pixazo_extract_audio
The source track, reused as the voice-clone reference
9 s
03b Clone the voice
pixazo_add_voice
A voice built from the clip’s own speech — no separate voice sample was supplied
~10 min (the bottleneck)
03c Speak
text-to-speech in the cloned voice
The translated line in that voice
42 s
04 Re-sync the mouth
pixazo_lip_sync, sync_mode “remap”
The picture time-fitted to the new audio
3 min 09 s
End to end ≈ 15 minutes · whole session $1.88, including generating the source clip · which vendor model performed the text-to-speech or the lip-sync was not disclosed by this run, so this page does not name one.
These two clips are the only genuine dub output on this page. Every other image and loop here is an illustration generated to explain a stage, and every one of them carries a visible “Illustration” label and says so in its alt text.
Four stations
How is the dub built, station by station?
Four stations run in order — transcribe, translate, voice, then lip sync — and each one hands its output straight to the next, which is why one instruction to the agent is enough to get a finished dub back. That chain is what separates an AI video language translator from a voiceover: the old picture is not left to drift, the mouth is rebuilt against the new spoken track.
Station 01 · Transcribe13 sIllustration
The speech becomes timed text
The agent listens to the clip you gave it and writes down what is actually said, with timings attached.
Tool
pixazo_transcribe
Returns
Text with timings, the detected language, and the speaker count
Station 02 · Translateno separate callIllustration
The agent writes the translation itself
There is no separate translation model in the chain — the agent phrases the new line directly.
Tool
the agent itself — no separate translation model
Returns
A line phrased to fit the spoken duration it measured
Station 03 · Voice~10 minIllustration
The speaker’s own voice is rebuilt, then speaks
The source track is pulled out, used as the clone reference, and then made to say the translated line.
The translated line in a voice cloned from that clip
the bottleneck
Station 04 · Re-sync3 min 09 sIllustration
The picture is time-fitted to the new audio
Remap re-times the footage against the translated take so the mouth lands where the new words do.
Tool
pixazo_lip_sync, sync_mode “remap”
Returns
The finished dub, mouth matched to the new line
What runs where
Where does AI VFX Studio stop and the playground begin?
AI VFX Studio is the only place that runs the whole chain: the playground has no speech-to-text and no translation model, so driving models there by hand covers the voice and the re-sync only, and you must bring your own translated script.
One route runs everything. The other runs two stations.
Give AI VFX Studio a clip and a target language and it walks all four stations itself. In the playground you are opening one model at a time, so the transcript and the translation have to arrive already written — there is nothing there that can produce either.
Lime square — AI VFX Studio runs it.
Lime outline — you supply it.
Crossed dark square — nobody does it.
The transcript, the translation, the cloned voice and the mouth re-sync all happen inside AI VFX Studio from a single instruction.
By hand in the playground you supply the translated script yourself, then run the voice station and the lip-sync station.
Neither route separates dialogue from music and effects, and neither replaces text that is burned into the picture.
Illustration
Illustration — generated to picture the re-sync stage, not dub output
Coverage by route. Every cell states its status in words as well as a symbol.
Job
AI VFX Studio (one instruction)
Playground, models driven by hand
Speech-to-text transcript
Covered. Covered
Not available. Not available — bring your own transcript
Translation
Covered. Covered — the agent writes it itself
Not available. Not available — no translation model here
Voice-clone reference
Covered. Covered — pulled from the clip’s own audio
Not verified. Not a verified capability here — the measured run never cloned a voice from a reference in the playground
Cloned-voice speech
Covered. Covered
Covered. Covered — Chatterbox v1, 4 credits
Lip re-sync to new audio
Covered. Covered
Covered. Covered — MultiTalk, from 104 credits
Separating dialogue from music and effects
Not available. Not available
Not available. Not available
Subtitle or SRT file out
Not available. Not available — the timed transcript is printed on screen, not exported as a subtitle file
Why does one five-second dub take about fifteen minutes?
Because cloning the voice is the bottleneck — roughly ten of those fifteen minutes go into building a voice from the clip’s own speech, and every other station finishes in seconds.
≈ 15 minutesend to end, one five-second single-speaker clip
Transcribe
13 s
Pull the source audio
9 s
Clone the voice — the bottleneck
~10 min
Speak the translated line
42 s
Re-sync the mouth
3 min 09 s
About 850 s of measured stage time$1.88 for the whole session, including generating the source clipOne run, one clip, one language pair
Bar widths are each station’s share of about 850 seconds of measured stage time, counting the roughly ten-minute voice clone as ten minutes flat; the translate step was not timed separately, so it is not in that total. This is one measured run on 29 July 2026, not a benchmark — a longer clip, more speakers or a different language pair will not produce the same numbers.
The voice, by hand
How do you generate just the spoken line yourself?
Open the Chatterbox v1 station in the playground at 4 credits, paste in the translated line you already have, and you get the spoken audio without running the rest of the chain. Used this way it is an AI voice translator for video rather than the full chain — the playground cannot transcribe or translate for you, so the script has to be yours. Your 100 signup credits cover this station many times over.
Illustration
The voice station on its own
One model, one line of text, one take — useful when the translation already exists and all you need is the speech.
Chatterbox v1 — Resemble AI · 4 credits
What it does. It speaks a line of text you provide, at 4 credits a run. It covers the same job as station 03 — speaking a line — as a model you run yourself, so you can iterate on a take without re-running a transcript or a re-sync. Which model ran inside the chain was never disclosed, so that is a comparison of jobs and not an identification.
What it will not do. It will not transcribe your clip and it will not translate anything — the playground has no speech-to-text and no translation model, so the translated script must arrive already written. It also does not touch the picture; the mouth is still saying the original words until you run a re-sync. Cloning a voice from a reference you upload is not something the measured run tested here, so this page does not claim this station does it.
Who approves the translation before any audio is generated?
You do, and it is the only checkpoint there is — the agent prints the transcript and the translated line on screen before it generates a single second of audio, because nobody has read-checked either and no accuracy figure has been published for the transcript or the translation.
The only checkpointIllustration
Illustration — the agent prints the transcript and the translated line on screen, and nothing is generated until you have read them.
Read the printed transcript
Check it against what is actually said in the picture. The transcriber is the first thing in the chain, so a word it mishears becomes a word the translation is built on.
Read the translated line
Check the meaning, and check it can be said inside the duration the transcriber measured. The agent phrases to fit that duration, which is a performance decision you are signing off on.
Then let it speak
Everything after this point inherits whatever you approved, in a cloned voice. The voice station and the re-sync do not re-read the script; they carry it forward exactly as approved.
If the line is wrong here, everything downstream is wrong — in the speaker’s own voice.
Before you approve, three things that are not the software’s call
Cloning a real person’s voice needs that person’s permission.
A licence to use footage does not extend to dubbing it, and has to be cleared separately.
The translated phrasing is fitted to the measured spoken duration, so it is a performance choice, not a literal gloss.
Who it’s for
Who is Pixazo’s AI video translator for?
Anyone who already owns finished footage in one language and needs the same speaker saying it in another. As an AI translator tool for video editing and video content teams, it fits course and tutorial teams, support and onboarding owners, product marketing, creator channels, internal comms, and localisation leads testing a language before commissioning a studio dub.
Course and tutorial teams
You start from a finished lesson recording with one instructor talking to camera, and you want the same instructor in a second language.
single speakerfront-facingtalking head
Support and onboarding owners
You start from a short walkthrough clip in your help centre where the voice matters more than the visuals.
under 60 sclean dialogueno burned-in titles
Product marketing
You start from an approved, locked cut and you need a second language without re-opening the edit.
finished cutlocked pictureone language pair
Creator channels
You start from your own to-camera footage, where the voice is yours to clone and permission is not a question.
own footageown voicesteady framing
Internal comms
You start from a leadership or policy message recorded once, and it has to reach teams who do not share that language.
single speakerseated shotscript on hand
Localisation leads
You start from one representative clip and want to hear a language before you commission a studio dub for the whole library.
short sampleone clipread the transcript
Holds / breaks
Which shots does the lip re-sync actually hold up on?
It holds up on a single speaker framed front-on and holding reasonably still, and it does not hold up on profile framing, on frames with two people talking, or under heavy motion blur — remap re-times the picture to the new audio, it does not repaint a face.
HoldsIllustration
One face, facing camera, mouth unobstructed — the re-sync has a clear reference to work from.
BreaksIllustration
In profile the mouth is barely in frame, and motion blur destroys what is left of the reference.
BreaksIllustration
With two speakers in one frame the re-sync has no way to know whose mouth to drive.
MultiTalk — from 104 credits
What it does. It re-syncs a mouth in existing footage to an audio track you hand it, from 104 credits a run. It covers the same job as station 04 — re-timing picture to new audio — as a model you run yourself, so you can re-time a clip against a take you produced elsewhere. Which model ran inside the chain was never disclosed, so that is a comparison of jobs and not an identification.
What it will not do. It will not translate, transcribe or produce the voice — the audio has to exist first. It also will not rescue framing it cannot read: profile shots, more than one speaker in frame, and heavy motion blur stay unreliable, and no setting changes that.
The Holds verdict is what the measured run showed, on one front-facing single-speaker clip. The two Breaks verdicts are known limits of remap re-timing, not results from this run — it contained no profile framing, no second speaker and no heavy motion blur. Nothing here is a scored test set, and there is no published pass rate.
Honest limits
What can’t an AI video translator do yet?
Four things flatly: it cannot pull dialogue back out of an already-mixed track, it cannot replace text burned into the picture, it cannot promise reliable lip-sync on profile framing, multi-speaker frames or heavy motion blur, and it cannot give you a guaranteed accuracy figure for the transcript or the translation.
Four things this page will not pretend to do.
Stems out of a finished mix
Dialogue, music and effects cannot be separated back out of an already-mixed track, so a dub replaces the whole audio bed.
Text burned into the picture
Titles, lower thirds and hard captions stay in the original language; they are pixels, not text.
Sync on the wrong framing
Profile shots, frames with more than one speaker, and heavy motion blur are unreliable, and no setting fixes that.
A guaranteed accuracy number
No accuracy figure has been published for the transcript or the translation, the on-screen print-out before audio generation is the only control, and if your workflow needs a contractual accuracy level this is not it.
What this run did not disclose
The agent named only Pixazo’s own tool names. The catalogue lists several speech and lip-sync families, but this run proves none of them, so this page does not claim which vendor model spoke or re-synced the clip above.
Before you dub footage you did not shoot.
What you need
Why it matters
Who has to agree
Permission to clone the speaker’s voice
A cloned voice is that person’s voice; consent to appear is not consent to be re-voiced
The speaker
A licence that covers redubbing
A licence to use a clip does not extend to putting new words in its speaker’s mouth
Whoever owns the footage
Sign-off on the translated line
The agent’s translation is fitted to the measured duration, so it is a phrasing choice, not a literal gloss
Whoever owns the message
Costs from the verified run: $1.88 for the whole session including generating the source clip; by hand, 4 credits at the voice station and from 104 credits at the lip-sync station. No per-minute or per-language pricing is published for a dub.
Four steps: upload a front-facing single-speaker clip, read the transcript and translated line the agent prints, approve them, then leave it alone while the voice clone runs — budget about fifteen minutes for a short clip.
Illustration
What the first run actually asks of you
One front-facing clip, one instruction, one read-through of what it prints — then about fifteen minutes of waiting, most of it the voice clone.
1
Upload the clip
One speaker, facing camera, mouth unobstructed. That framing is what the re-sync stage needs later, so it is worth getting right before anything else runs.
2
Say what language you want
One instruction is enough; the agent chooses the stations. You are not picking models or wiring a chain together yourself.
3
Read what it prints, then approve
The transcript and the translated line appear before any audio exists. This is your only checkpoint — read both properly.
4
Wait out the voice clone
Roughly ten minutes on a short clip, unattended; that pause is expected, not a stall. The re-sync that follows took 3 min 09 s in the measured run.
What are the most frequently asked questions about AI video translation?
These ten cover timing, voice samples, language pairs and direction, what a dub leaves untouched, cost and the free tier, permission, using a YouTube upload as the source, and what the run did not disclose.
How long does one dub take?
About fifteen minutes end to end for a five-second, single-speaker clip in the measured run on 29 July 2026, and roughly ten of those minutes were the voice clone alone.
Transcribing took 13 seconds, pulling the source audio 9 seconds, speaking the translated line 42 seconds and re-syncing the mouth 3 minutes 9 seconds. A longer clip, more speakers or a different language pair will not produce the same numbers.
Do I need to upload a separate voice sample?
No — the run cloned the voice from the clip’s own speech, with no separate sample supplied.
The source track is pulled out with pixazo_extract_audio and reused as the clone reference, which also means a noisy or music-heavy original gives the clone a worse reference to work from.
Which language pairs work?
The agent writes the translation itself rather than calling a separate translation model, and the verified run was English to Spanish.
Because there is no separate translation model in the chain, there is no published list of supported pairs to quote, and no accuracy figure for the translation either. Note also that the playground has no speech-to-text and no translation model at all — only AI VFX Studio runs the whole chain.
Does a dub touch anything besides the dialogue?
It replaces the whole audio bed and re-times the picture, and it touches nothing else.
It cannot separate dialogue from music and effects inside an already-mixed track, and it cannot replace text burned into the picture — titles, lower thirds and hard captions stay in the original language.
What does it cost?
The whole verified session came to about $1.88, and that figure includes generating the source clip, not just the dub.
If you drive the stations by hand instead, Chatterbox v1 is 4 credits and MultiTalk is from 104 credits per run. There is no published per-minute or per-language dub price.
Do I need permission to clone someone’s voice?
Yes — cloning a real person’s voice needs that person’s permission, and a licence to use footage does not extend to dubbing it.
Clear both separately before you dub anything you did not shoot yourself.
Which vendor model spoke and re-synced the clip?
Not disclosed: the agent named only Pixazo’s own tool names — pixazo_add_voice, pixazo_lip_sync.
The catalogue lists several speech and lip-sync families, but this run proves none of them, so this page does not attribute either step to a vendor.
Can it translate video to English, or only out of English?
Either direction. The transcribe stage detects the source language itself — in the measured run it reported “Detected English” without being told — so you name the language you want out and it works from whatever went in.
The run on this page went English to Spanish. Spanish to English, or any other pair, is the same instruction with a different target language; coverage of the target is set by the speech model that speaks the line, not by Pixazo.
Can I translate a YouTube video with it?
You need the video file, not a link. Nothing in the measured run imported from a URL, so export the clip from YouTube — or use your own master — and upload that.
A single speaker facing camera is what the re-sync stage handles best, so a talking-head upload dubs far more cleanly than a fast-cut montage.
Is there a free AI video translator tier?
Yes, free to start: a 7-day trial with no card required and 100 bonus credits, which is what the pricing page states. Those credits go a long way at the voice station — Chatterbox v1 costs 4 credits a line.
Be realistic about the whole chain, though. The lip re-sync station starts at 104 credits, so a complete end-to-end run needs more than the signup grant. The measured run on this page cost about $1.88 in AI VFX Studio, and current plan details are on the Pixazo pricing page.
DJ
Deepak Joshi
Content Marketing Specialist · Pixazo
Deepak documents Pixazo’s speech, voice-cloning and lip-sync models by
running them rather than reading their spec sheets. The dub on this page is his own test run in
AI VFX Studio on 29 July 2026 — the transcript, the translation, the step timings and the
$1.88 session cost are all copied from that run, and the two clips are its actual output.
Where the run did not reveal something, such as which vendor model spoke the line, this page
says so instead of guessing.
What other Pixazo tools pair with the AI video translator?
The ones that make or extend the footage you are dubbing — start upstream if you still need the video, and use the directory if you are looking for a different job entirely.
Drop one front-facing, single-speaker clip into AI VFX Studio and approve the transcript and translation it prints — that is exactly how the two clips at the top of this page were made.