The first thing I heard was my own sentence coming out of a stranger's mouth.
I'd written the script in Arabic, read it in my head maybe forty times in my own voice, and then a deep, unhurried Saudi narrator said it back to me with a weight I hadn't put there. Every word landed like it mattered. It was genuinely good audio. It was also wrong for the video, and it took me two more full renders to work out why.
That's the part nobody tells you: choosing an AI voice is casting. Same decision a director makes when three people read for the same part, except it costs cents instead of a studio day. I treated it like a dropdown right up until I heard three takes back to back and realized I'd been about to publish the wrong one.
Why I needed a voice that doesn't speak English
This site is my English project. I also run a second one in Arabic — Thukaa, a daily AI-news outlet with a YouTube channel at @thukaa_ai. It needs narration in Arabic every single day, which has parked me in a corner of AI voice that English-language blogs never write about, because they've never had to.
The job here was a vertical YouTube Short: 40.6 seconds, four clips, Vox-style motion graphics. The argument is that the AI bubble is money circulating in a loop — the same dollars going round between the same handful of companies and getting counted as growth each time around. Punchy, a bit sardonic, needs to move fast.
The visuals came from elsewhere — stills out of Nano Banana Pro, motion out of Seedance 2.0 — and I'll write that pipeline up another time. This is about the voice, because the voice is what made or broke it.
Three takes, one script
I used the ElevenLabs web interface with the Multilingual v2 model. Finding Arabic voices is unglamorous: you type "arabic" into the voice picker and read the names people gave their uploads. My first pick was called Ali — calm & Deep Arabic Saudi Narrator, which as a Saudi felt like the obvious call. Matching accent, matching audience, done.
Then I did the thing I'd recommend to anyone: instead of guessing from a five-second preview, I rendered the entire narration three times, with three different voices. Ali. Hazem. Mamdoh. Same script, same settings, three complete takes, played back to back like demos.
You can't hear them from where you're sitting, so here's what separated them.
Ali was the nature documentary. Deep, slow, enormous authority, settling downward at the end of every sentence like he was closing a door behind each thought. Beautiful for a three-minute explainer. Across forty seconds and four hard cuts, that gravity turned into drag — he was still finishing a sentence while the visual had already moved on.
Hazem solved the pace problem and created a new one. Brighter, faster, newsreader-forward, pushing through the lines with real momentum. But he gave every clause the same amount of energy. The setup and the punchline came out at identical weight, so the joke in the middle of the script just... went past. A flat read at speed is still a flat read.
Mamdoh was Egyptian-accented, mid-weight, and conversational in a way the other two weren't. Where Ali settled and Hazem pushed, Mamdoh leaned — he'd put a little extra pressure on the operative word in a clause and lift slightly going into the last one, which is exactly the shape you want on a line about money going around in a circle. He sounded like someone explaining something to you, not reciting it at you.
The one that shouldn't have won
Mamdoh won. I'm Saudi, my audience skews Gulf, and I picked the Egyptian voice.
On paper that's the wrong answer. Accent-matching is the first rule everybody repeats, and Ali was built for exactly my slot. But the script wasn't a documentary, it was a wry forty-second argument, and Mamdoh read it better — not marginally, obviously, once the three were next to each other.
Accent tells you where a voice is from. It doesn't tell you whether it can land your joke.
Worth noting alongside that: ElevenLabs' docs list Multilingual v2 as covering 29 languages, and the Arabic entry reads "Arabic (Saudi Arabia, UAE)." Egyptian isn't on the list at all — because that list describes the model's language coverage, while the accent you actually hear belongs to the individual voice, not to a setting you flip. Which means you can't shop for these off the spec sheet. You have to listen.
The bug nobody warns you about in English
Here's the thing I'd have paid money to know beforehand.
I wrote the script the way I write everything, with numbers as numerals. The render came back and the numbers were wrong — not slightly accented, wrong, in a way that would make a viewer stop the video.
The reason is obvious in hindsight and invisible until it bites you. A numeral is a shape, not a word. "2" carries no information about whether to say two, deux, or ithnayn. The model has to guess a language for that shape, and outside English it guesses badly. ElevenLabs' own best-practices page arrives at the same place from the other end: normalize the text yourself, turn "$1,000,000" into "one million dollars" before you send it, because models interpret numeric formats inconsistently. The multilingual model handles it better than the small fast one, and "better" is not "reliably."
So I went back through the script and rewrote every digit as an Arabic word, spelled out exactly as I wanted it said. Re-rendered. Fixed.
Generalize it and it's the most useful thing in this article: in any non-English TTS, never send a digit. Spell out numbers, dates, currencies and acronyms as words in the target language. Then do the second half of the rule — read your script out loud in that language before you render anything. Your mouth catches the ambiguity your eyes skip over.
The part that genuinely sucked
The number fixes weren't a one-shot repair, and that's where the day went sideways.
I was rendering per-segment — one audio file per section of the video, because that's what lets you cut precisely. So every number I fixed meant re-rendering that segment, which produced a file of a slightly different length, which meant everything after it started at a different timestamp. Small edits, cascading downstream. Meanwhile the browser hands you a stack of downloads with auto-generated filenames that all look identical by the eleventh one, so a real fraction of my afternoon went on playing three-second clips to work out which take I was holding.
None of it is hard. It's fiddly in the way that eats an hour and produces nothing you can show anyone. Doing it again, I'd rename every file the second it lands, segment number first.
Record first, cut second
The other decision that saved me: I built this audio-first.
The obvious way round is to generate the video, then lay a voice over it. Don't. Video generators hand you clips of whatever length they feel like, and forcing narration onto a fixed picture means either rushing the read or padding it with dead air. Both sound exactly like what they are.
So the narration got locked first — per segment, silence trimmed off the head and tail of each file — and each clip's duration was then cut to match its segment of audio. The picture serves the voice. When a clip came back shorter than the audio it had to cover, I froze its last frame to close the gap:
# hold the final frame for 1.4s so the clip covers its narration
ffmpeg -i clip3.mp4 \
-vf "tpad=stop_mode=clone:stop_duration=1.4" \
-c:a copy clip3_padded.mp4
A held frame under a continuing sentence reads as deliberate. A sentence chopped off to fit a clip reads as broken. Nobody watching a Short is auditing your frame counts, but everybody hears bad timing.
What three auditions actually cost
Almost nothing, which is the whole argument for doing it.
The tier I was working on gives 10,000 credits a month. After three complete versions of the narration plus every number-fix re-render, I still had roughly 9,000 left — three full auditions of a 40-second script came to about a tenth of one month's allowance. Rendering the whole thing three times isn't extravagant. It's the cheapest quality decision in the entire pipeline.
When you shouldn't use one of these at all
I'll happily hand Thukaa's daily read to a synthetic voice. Thukaa is an outlet — the audience's relationship is with the reporting and the fact that it shows up every day, not with my larynx.
But there are places I won't take one, and the line is clearer than the discourse makes it:
- Anywhere the relationship is with you personally. A founder update, a channel built on your face, a teaching video where trust is the product. If people subscribed to you, a synthetic read is a small betrayal even when it's technically excellent.
- Anything with real people's names in it. Interviews, credits, tributes. TTS mispronounces names in every language, and getting someone's name wrong in public is a real cost, not a rendering artifact.
- Emotion you haven't earned. Apologies, condolences, gratitude. A machine performing sincerity is worse than silence.
- Anything you'd be embarrassed to label. That's the whole test. If you couldn't comfortably write "narration is AI" in the description, you already know the answer.
This site is written in first person, and if I ever publish a Vibe Gate video, it'll be my actual voice, badly recorded, with the room tone in it. The Arabic news read is a job. This isn't.
Go listen
The Short is public, so you don't have to take my word for any of it: youtube.com/shorts/FhV_cMTIPE8. Forty-point-six seconds, four clips, the AI bubble going round in a circle, narrated by an Egyptian voice that beat the one that was supposed to win.
If you take one habit from this: render the whole script in more than one voice before you commit. Not the preview — the whole thing. It costs about as much as nothing, and it's the only way you'll find out that the obvious pick was the wrong pick.
Sources
Everything checked July 25, 2026:
- ElevenLabs — Pricing (Free tier: 10,000 credits/month, no commercial licence, attribution required; Starter: $6/month, 30,000 credits, commercial licence)
- ElevenLabs Docs — Models ("Our multilingual v2 models support 29 languages", including Arabic (Saudi Arabia, UAE))
- ElevenLabs Docs — Text to Speech best practices (normalizing numeric formats before generation; Multilingual v2 vs Flash v2.5 on number handling)
- ElevenLabs Help Center — Why are numbers, dates, symbols and acronyms not properly pronounced or spoken in the correct language? (write numbers out in words in the target language)
- The finished Short — Thukaa, published on YouTube
Previously on the build-log beat: how this site got made in a terminal. Next up, the visual half of this pipeline — stills, motion, and why four clips was the right number.