Synthesizing high-fidelity voice output from text that contains mixed-language scripting (e.g., Arabic-English code-switching) is one of the most difficult challenges in generative speech science. When a voice engine transitions between different language systems, it must preserve natural cadence, speaking style, and most importantly, speaker identity.
From Ragini to Rawi: Evolution of Synthesis
At ConvoZen Research, our journey into multilingual generative speech began with Ragini—our foundational voice synthesis engine developed for complex Indian language code-switched environments like Hindi-English and Tamil-English. Ragini pioneered the decoupling of speaker embedding vectors from phonemic language representations, proving that cross-lingual voice matching is possible without identity distortion.
Building directly upon the structural breakthroughs of Ragini, we developed Rawi v1 to solve the intricate acoustic challenges of Arabic speech—combining native code-switching with rich multi-dialect conditioning.
Unrivaled Fine-Grained Dialectal Conditioning
Beyond bilingual fluidity, real-world enterprise deployment across the Middle East demands authentic regional accent conditioning. Most global commercial TTS engines fail completely when exposed to regional speech, defaulting to rigid, synthetic Modern Standard Arabic (MSA) inflections that sound artificial to local ears.
Rawi v1 stands as one of the few generative speech models in the global AI landscape to achieve such fine-grained dialectal control natively—capturing distinct sub-dialect micro-variations across Najdi, Hijazi, Emirati, Egyptian, and Jordanian speech within a single unified neural architecture.
Rawi v1 embeds this granular conditioning natively across four core speech pillars:
Comprehensive coverage for Najdi, Hijazi, and Emirati colloquial cadence and vocabulary blend.
Natural phonology matching Cairene and regional Egyptian speech dynamics combined with technical English loanwords.
Precise pitch intonation and vowel length modeling for Jordanian Levantine conversational speech.
Pristine broadcast-quality MSA synthesis for formal news, documentation, and educational media.
Ultra-Low Latency Streaming (RTF & TTFB)
To power interactive voice agents and real-time telephony, raw speech quality is only half the battle—latency and hardware throughput are paramount. Under real-world deployment benchmarks on standard commodity hardware (NVIDIA T4 GPU), Rawi v1 achieves an efficient Real-Time Factor (RTF 0.47) and a rapid Time-To-First-Audio-Buffer (TTFB 260ms).
By generating and streaming audio chunks in continuous neural buffers, Rawi allows AI agents to speak back to users instantaneously, maintaining natural human-like turn-taking without awkward pauses.
Select a dialect sample to listen to generated voice output:
"Go ahead and spell out the new street name and provide the building number so I can update your service address."
"لَا تَقْلَقْ، سَنُعِيدُ حِسَابَ الْوَزْنِ الْحَجْمِيِّ لَكَ وَنُدَقِّقُ الْقِيَاسَ لِضَمَانِ صِحَّتِهِ."
"أَهَا، تِبْغِي مُرَاجَعَةْ تَسْهِيلْ اَلسَّحْبْ عَلَى اَلْمَكْشُوفْ لِلـ account؟"
"رَاجْعِي شُرُوطْ المِنْحَةْ الدِّرَاسِيَّةْ بِالمَوْقِعْ عَشَانْ تِعْرِفِينْ كُلْ التَّفَاصِيلْ المَطْلُوبَةْ لِلتَّقْدِيمْ."
"حَالِيًّا الفَرِيقْ المُخْتَصْ بيِشْتَغِلْ عَلَى طَلَبِكْ، وَأَوَّلْ مَا يِجِينِي مِنْهُمْ أَيّ رَدْ أَوْ تَحْدِيثْ، حَأَكُونْ أَنَا أَوَّلْ وَحْدَةْ أَتْوَاصَلْ مَعَاكِ وأَطَمِّنِكْ."
"هَلْ حَضْرِتِكْ مُحْتَاجَة أَيْ مُسَاعَدَة فِي خَطْوَة تَانْيَة؟"
"وَصَلْنِي رَمْزْ التَّحَقُّقْ تَبَعَكْ، وَتَمْ تَأْكِيدْ هُوِيَّتَكْ بِنَجَاحْ."
Scientific Evaluation Framework
Evaluating generative speech science requires replacing subjective listening tests with an objective, standardized benchmarking suite. We evaluate synthesized speech across five core scientific dimensions: Intelligibility, Perceptual Naturalness, Speaker Identity Consistency, Prosody Dynamics, and Spectral Purity.
Intelligibility (ASR)
Measures transcription accuracy at the word, character, and phoneme level. Lower is better.
Proxy for model confidence using NeMo ASR. Higher is better.
Naturalness & Quality
Neural estimates of overall audio quality, signal purity, and background cleanliness on a 1-5 scale.
Mathematical predictions of perceptual clarity (PESQ/STOI) and signal-to-distortion ratio (SI-SDR).
Speaker Consistency
Measures voice stability and similarity across sliding windows. Ensures the speaker identity doesn't fluctuate.
Similarity between start and end of clips to catch any gradual voice distortion over time.
Prosody & Dynamics
Pitch inflection spread. Higher indicates more expressive speech, lower indicates monotone.
Measures spectrum flatness and high-frequency ratio to prevent metallic or buzzy audio artifacts.
Empirical Benchmark Evaluation
We thoroughly benchmarked Rawi v1 against global industry leaders (including ElevenLabs, Google Chirp, and Cartesia) using this objective scientific framework.
Empirical Performance Highlights & Competitive Leadership
Rawi v1 leads all evaluated engines with an ASR confidence score of -49.3 (outperforming ElevenLabs -50.1, Cartesia -52.2, and Google Chirp -53.8), demonstrating unmatched phonetic stability.
With 0.366 WER and 0.226 CER, Rawi v1 significantly outperforms both ElevenLabs (0.402 WER / 0.263 CER) and Cartesia (0.388 WER / 0.264 CER) on bilingual code-switched text.
Achieves a speaker consistency score of 0.678, outperforming ElevenLabs (0.640) and demonstrating steady voice identity across code-switched audio streams.
Scores 3.27 OVRL and 4.12 BAK, outperforming ElevenLabs (3.16 OVRL / 3.98 BAK) and Cartesia (3.25 OVRL / 4.03 BAK) while proving highly comparable to Google Chirp (3.41).
| Model | WER ↓ | CER ↓ | PER ↓ | TTScore-int ↑ |
|---|---|---|---|---|
| Rawi v1 (Ours) | 0.366 | 0.226 | 0.385 | -49.3 |
| Google Chirp | 0.309 | 0.153 | 0.304 | -53.8 |
| Cartesia | 0.388 | 0.264 | 0.438 | -52.2 |
| ElevenLabs | 0.402 | 0.263 | 0.438 | -50.1 |
| Model | DNSMOS-OVRL ↑ | DNSMOS-SIG ↑ | DNSMOS-BAK ↑ | SQUIM-PESQ ↑ | SQUIM-STOI ↑ |
|---|---|---|---|---|---|
| Rawi v1 (Ours) | 3.27 | 3.52 | 4.12 | 2.72 | 0.971 |
| Google Chirp | 3.41 | 3.63 | 4.16 | 3.68 | 0.993 |
| Cartesia | 3.25 | 3.53 | 4.03 | 4.08 | 0.995 |
| ElevenLabs | 3.16 | 3.46 | 3.98 | 3.42 | 0.991 |
| Model | SpkConsist(mean) ↑ |
|---|---|
| Rawi v1 (Ours) | 0.678 |
| Google Chirp | 0.701 |
| ElevenLabs | 0.640 |
| Cartesia | 0.687 |
Across comprehensive empirical evaluations, Rawi v1 establishes robust, high-fidelity voice synthesis performance for bilingual Arabic-English speech—delivering clear round-trip transcription intelligibility, steadfast speaker identity preservation, and responsive real-time streaming capability.
Try out Rawi v1 from the ConvoZen platform. For any inquiries or to request early access to our next-generation models, reach out to the research team at contact@convozen.ai.